arXiv cs.CL
· Papers
Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs
arXiv:2601.21476v2 Announce Type: replace Abstract: On-policy reinforcement learning (RL) for language model post-training suffers from a fundamental tension: as training progresses, policy entropy collapses and sampling diversity diminishes, causing the model to ``forget'' its own earlier exploratory capacity. While o