arXiv cs.CL
· Papers
Reasoning Fine-Tuning Induces Persistent Latent Policy States
arXiv:2607.18532v1 Announce Type: new Abstract: Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood. It is unclear whether reasoning fine-tuning improves local token-level competence or globa