arXiv cs.LG
· Papers
Towards Isolated Interventions via Almost Orthogonal Features in Language Models
arXiv:2602.04718v2 Announce Type: replace Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substantially alter the ef