Skip to content
arXiv cs.CV · Papers

Logit Lens Supervision for Patch-Level Explanations in Vision-Language Models

arXiv:2602.01530v2 Announce Type: replace Abstract: Modern autoregressive Vision-Language Models (VLMs) can generate fluent answers while their visual-token representations become weakly tied to the image regions from which they originate. This limits patch-level explainability: a visual token should remain interpretab