arXiv cs.CV
· Papers
TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions
arXiv:2607.26107v1 Announce Type: new Abstract: Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for these tasks by learning