Trust Region Masking for Long-Horizon LLM Reinforcement Learning
arXiv:2512.23075v5 Announce Type: replace-cross Abstract: Policy gradient methods for Large Language Models optimize a policy $pi_theta$ via a surrogate objective computed from samples of a rollout…