Skip to content
arXiv cs.CL · Papers

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models

arXiv:2607.07251v1 Announce Type: new Abstract: One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, which are defined as spatial expressions w