Skip to content
LessWrong AI · Communities

Evidence for feature-specific error correction in LLMs

LLMs are commonly assumed to use superposition to represent more features than they have dimensions. The evidence for this is mostly indirect — chiefly the success of SAEs at extracting interpretable directions. A stronger claim is that models also compute in superposition, and for that we have only theoretical evidenc