arXiv cs.AI
· Papers
Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
arXiv:2607.19364v1 Announce Type: new Abstract: Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering methods often rely on learned steering objectives or single-criterion feature selection. We introduce a transparent SAE-feature stee