Skip to content
arXiv cs.CL · Papers

Revealing Hidden Model Behaviors with Task-Specific Self-Reports

arXiv:2607.03640v2 Announce Type: replace Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic. We introduce the Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter tha