Skip to content
arXiv cs.LG · Papers

Towards Isolated Interventions via Almost Orthogonal Features in Language Models

arXiv:2602.04718v2 Announce Type: replace Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substantially alter the ef