Skip to content
arXiv cs.CV · Papers

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation

arXiv:2608.02474v1 Announce Type: new Abstract: Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its inference remains expensive due to the iterative denoising process of diffusion models. Existing caching methods mainly explo