arXiv cs.CL
· Papers
Revealing Hidden Model Behaviors with Task-Specific Self-Reports
arXiv:2607.03640v2 Announce Type: replace Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic. We introduce the Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter tha