Skip to content
arXiv cs.CL · Papers

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

arXiv:2607.24645v1 Announce Type: cross Abstract: The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts