Skip to content
arXiv cs.CL · Papers

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the information a defender most needs, namely how harmful the resulting action wa