Skip to content
LessWrong AI · Communities

Is there even a ground-truth for LLMs’ internal representations?

[This is an introductory blog for the paper Laguerre Geometry for Interpreting Large Language Models and the GitHub repository Geometric Lens.]LLM Lens: What does an internal vector mean?Anthropic's recent paper on the "J-Lens" (Jacobian Lens) has revived interest in reading the "thoughts" inside Large Language Models.