Skip to content
arXiv cs.LG · Papers

Mask-Aware Policy Gradients for Diffusion Language Models

arXiv:2607.15200v1 Announce Type: cross Abstract: Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate thi