Skip to content
arXiv cs.CV · Papers

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from detailed textual descriptions. While methods such as Long-CLIP extend the token limit through positional embedding interp