arXiv cs.CL
· Papers
H$^2$SD: Hybrid Hindsight Self-Distillation
arXiv:2607.18955v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level guidance. Existing self-distillation methods add a privileged teacher but typicall