Skip to content
arXiv cs.CL · Papers

H$^2$SD: Hybrid Hindsight Self-Distillation

arXiv:2607.18955v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level guidance. Existing self-distillation methods add a privileged teacher but typicall