arXiv cs.CL
· Papers
DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech
arXiv:2607.04140v2 Announce Type: replace-cross Abstract: Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness, since local errors propagate to later positions and can escalate into hallucination. This limitation stems from t