Skip to content
arXiv cs.CL · Papers

Length Penalties Make Chain-of-Thought Less Monitorable

arXiv:2607.09786v2 Announce Type: replace-cross Abstract: Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints from steering models, even though the model