Skip to content
arXiv cs.AI · Papers

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models

arXiv:2607.16339v2 Announce Type: replace Abstract: Diffusion-based Large Language Models(DLLMs) enable parallel generation via Semi-Autoregressive (SAR) decoding in text generation. However, current methods suffer from severe operator-level redundancy: they recompute the entire sequence during denoising steps, ignorin