Skip to content
HF Daily Papers · Papers

dOPSD: On-Policy Self-Distillation for Diffusion Language Models

Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains difficult: supervised fine-tuning is off-policy and suffers from exposure bias, whi