Skip to content
arXiv stat.ML · Papers

Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers

arXiv:2606.11949v3 Announce Type: replace-cross Abstract: Reasoning models deployed as safety monitors exhibit a systematic vulnerability: reasoning-token budget starvation. Adversarial inputs require $3.3times$ more reasoning tokens than benign inputs to produce valid safety scores ($T_{50,text{adv}}{=}154$ vs. $T_{