arXiv cs.CV
· Papers
Logit Lens Supervision for Patch-Level Explanations in Vision-Language Models
arXiv:2602.01530v2 Announce Type: replace Abstract: Modern autoregressive Vision-Language Models (VLMs) can generate fluent answers while their visual-token representations become weakly tied to the image regions from which they originate. This limits patch-level explainability: a visual token should remain interpretab