arXiv cs.CL
· Papers
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models
arXiv:2607.07251v1 Announce Type: new Abstract: One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, which are defined as spatial expressions w