arXiv cs.CV
· Papers
Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment
arXiv:2607.04311v2 Announce Type: replace Abstract: Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity consistency and model complex relationships among multiple subjects. In this paper, we propose Aura, a unified frame