arXiv cs.CV
· Papers
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
arXiv:2607.23855v2 Announce Type: replace-cross Abstract: Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains chal