LessWrong AI
· Communities
The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
TL;DRCurrent model organisms (MOs) for interpretability benchmarking are typically constructed via a dedicated, “post-hoc” SFT step. However, recent work suggests that this may make interpretability unrealistically easy, giving the field misplaced confidence in the readiness of interpretability techniques to audit safe