Skip to content
arXiv cs.CL · Papers

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech

arXiv:2607.04140v2 Announce Type: replace-cross Abstract: Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness, since local errors propagate to later positions and can escalate into hallucination. This limitation stems from t