PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR
arXiv:2608.11368v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories. Recent allocators reduce this…