arXiv cs.CV
· Papers
BabyVision: Visual Reasoning Beyond Language
arXiv:2601.06521v2 Announce Type: replace Abstract: While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently