Skip to content
arXiv cs.CV · Papers

Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task

arXiv:2607.01503v1 Announce Type: new Abstract: In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model performance. To this end, we combine depth-ordering and odd-one-out psychophysical tasks: the V