arXiv cs.CV
· Papers
Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
arXiv:2608.04124v1 Announce Type: new Abstract: Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time. Existing methods typically rely on long textual chain-of-thought rationales, even though many questions can be answered as so