arXiv cs.CV
· Papers
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
arXiv:2602.00462v5 Announce Type: replace Abstract: Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM. Intriguingly, this mapping can be as simple as a shallow MLP transformation. To understa