LessWrong AI
· Communities
Red-teaming LLM unlearning: LUNAR's "forgotten" knowledge is still recoverable
Summary LUNAR is a state-of-the-art unlearning method. To forget specific (harmful) knowledge, it retrains a single MLP down-projection matrix such that activations from this “forget” set are redirected into regions that produce “I don’t know” responses. Under standard evaluation LUNAR looks robust, including against w