Skip to content
arXiv cs.CL · Papers

Reasoning Fine-Tuning Induces Persistent Latent Policy States

arXiv:2607.18532v1 Announce Type: new Abstract: Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood. It is unclear whether reasoning fine-tuning improves local token-level competence or globa