Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models
arXiv:2607.07251v1 Announce Type: new Abstract: One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To…