Skip to content
arXiv cs.LG · Papers

Can Interpretation Predict Behavior on Unseen Data?

arXiv:2507.06445v3 Announce Type: replace Abstract: Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this alternate objective by using model internals to predict their out-of-distribution (OOD) b