r/MachineLearning
· Communities
Agentic safety triggers aren’t textual safety triggers — MCP attacks that beat SOTA guardrails more than half the time (code + dataset) [R]
Most safety alignment work treats "detect the attack" as a text classification problem — does the prompt contain language the model's safety guardrails should catch. That assumption breaks down for LLM agents with real tool access. Here's a concrete case: take a known, public security vulnerability (a CVE), work out th