Skip to content
arXiv cs.LG · Papers

Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance

arXiv:2607.19386v1 Announce Type: new Abstract: Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For these comparisons to be meaningful, scores