arXiv cs.CV
· Papers
EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation
arXiv:2608.02474v1 Announce Type: new Abstract: Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its inference remains expensive due to the iterative denoising process of diffusion models. Existing caching methods mainly explo