Skip to content
LessWrong AI · Communities

Interpretability is becoming increasingly uninterpretable

What is the purpose of interpretability research? Anthropic states that the mission of their interpretability team is to "discover and understand how large language models work internally, as a foundation for AI safety and positive outcomes". I think this characterization constitutes the classical argument for studying