arXiv cs.CV
· Papers
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
arXiv:2607.20868v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturally develop through continuous observation of the real world, such as spatial perception and d