Skip to content
arXiv cs.CL · Papers

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

arXiv:2505.14300v2 Announce Type: replace-cross Abstract: White-box monitoring is increasingly adopted as an auditing tool as Large Language Models (LLMs) are deployed in daily operations to ensure safe model behavior. However, white-box monitors can be circumvented, and the mechanisms underlying such evasion have not