arXiv cs.AI
· Papers
LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models
arXiv:2607.16339v2 Announce Type: replace Abstract: Diffusion-based Large Language Models(DLLMs) enable parallel generation via Semi-Autoregressive (SAR) decoding in text generation. However, current methods suffer from severe operator-level redundancy: they recompute the entire sequence during denoising steps, ignorin