HF Daily Papers
· Papers
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages exploration beyond it. Their c