Skip to content
arXiv cs.CL · Papers

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

arXiv:2608.00011v1 Announce Type: new Abstract: Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and decode tokens sequentially, while non-autoregressive approaches im- prove speed at the cost of l