Skip to content
arXiv cs.CL · Papers

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

arXiv:2608.07763v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which limits their ability to handle culturally grou