Skip to content
arXiv cs.CL · Papers

Weak-to-Strong Generalization via Direct On-Policy Distillation

arXiv:2607.05394v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because the target model must generate many rollouts during training. As models scale, post