Skip to content
arXiv cs.CV · Papers

Video Generation Models are General-Purpose Vision Learners

arXiv:2607.09024v1 Announce Type: new Abstract: Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video