Skip to content
LessWrong AI · Communities

Refusal Is Redundantly Distributed, Not Localized: A Per-Layer Ablation Study on Llama-3.1-8B

TL;DRThis work replicates and extends the findings of Arditi et al. [1], who studied the refusal mechanism and found that a single direction, obtained through Difference-in-Means (DIM) methods, is enough to causally ablate and steer the model behavior.The project builds on those results through two additional experimen