Skip to content
arXiv cs.AI · Papers

Trust Region Masking for Long-Horizon LLM Reinforcement Learning

arXiv:2512.23075v5 Announce Type: replace-cross Abstract: Policy gradient methods for Large Language Models optimize a policy $pi_theta$ via a surrogate objective computed from samples of a rollout policy $pi_{text{roll}}$. However, modern LLM-RL pipelines suffer from unavoidable implementation divergences -- backe