arXiv cs.CV
· Papers
A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval
arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from detailed textual descriptions. While methods such as Long-CLIP extend the token limit through positional embedding interp