arXiv cs.CL
· Papers
Length Penalties Make Chain-of-Thought Less Monitorable
arXiv:2607.09786v2 Announce Type: replace-cross Abstract: Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints from steering models, even though the model