Skip to content
arXiv stat.ML · Papers

When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers

arXiv:2608.12623v1 Announce Type: cross Abstract: Language model classifiers with explanations are used for moderation, routing, topic triage, and low-resource annotation. We study black-box auditing when the defender has only clean calibration data without trigger information but can ask the classifier for a label plu