Skip to content
arXiv cs.AI · Papers

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches

arXiv:2607.16205v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrasting multiple self generated rollouts. However, we identify a critical support limited bottle