arXiv cs.CV
· Papers
Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task
arXiv:2607.01503v1 Announce Type: new Abstract: In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model performance. To this end, we combine depth-ordering and odd-one-out psychophysical tasks: the V