Skip to content
arXiv cs.LG · Papers

Activation Probes Surface Code-Security Signals that the Model's Output Misses

arXiv:2608.09643v1 Announce Type: cross Abstract: AI coding agents now write a growing share of production code, and human security review does not scale at the rate code is generated. The agents in widest use are closed-weight, so a deploying team cannot read their internals. It can instead run an open-weight model as