arXiv cs.AI
· Papers
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
arXiv:2607.16205v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrasting multiple self generated rollouts. However, we identify a critical support limited bottle