arXiv cs.AI
· Papers
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
arXiv:2512.23075v5 Announce Type: replace-cross Abstract: Policy gradient methods for Large Language Models optimize a policy $pi_theta$ via a surrogate objective computed from samples of a rollout policy $pi_{text{roll}}$. However, modern LLM-RL pipelines suffer from unavoidable implementation divergences -- backe