Skip to content
arXiv cs.CV · Papers

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs

arXiv:2602.00462v5 Announce Type: replace Abstract: Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM. Intriguingly, this mapping can be as simple as a shallow MLP transformation. To understa