arXiv cs.CL
· Papers
Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the information a defender most needs, namely how harmful the resulting action wa