Skip to content
arXiv cs.AI · Papers

Robust Critics: Defending LLMs Against Multi-Turn Attacks

arXiv:2607.20472v1 Announce Type: new Abstract: When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the central challenges of LLM safety. A model that assumes the worst harms legitimate users; one that assumes the best is eas