Skip to content
LessWrong AI · Communities

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

TL;DRCurrent model organisms (MOs) for interpretability benchmarking are typically constructed via a dedicated, “post-hoc” SFT step. However, recent work suggests that this may make interpretability unrealistically easy, giving the field misplaced confidence in the readiness of interpretability techniques to audit safe