Skip to content
LessWrong AI · Communities

Red-teaming LLM unlearning: LUNAR's "forgotten" knowledge is still recoverable

Summary LUNAR is a state-of-the-art unlearning method. To forget specific (harmful) knowledge, it retrains a single MLP down-projection matrix such that activations from this “forget” set are redirected into regions that produce “I don’t know” responses. Under standard evaluation LUNAR looks robust, including against w