Skip to content
arXiv cs.AI · Papers

Adaptively Robust LLM Monitoring via Activation Watermarking

arXiv:2603.23171v3 Announce Type: replace-cross Abstract: Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent. LLM monitoring is deterministic and often openly available, so $emph{adaptive}$ attackers with a local copy can search offline for prompts that elicit harmful beh