LessWrong AI
· Communities
Confirming Claims of Superposition and Adversarial Examples in Toy Models
This is a replication of Adversarial Attacks Leverage Interference Between Features in Superposition, completed as part of the Second Look Summer Fellowship.tl;dr:We reproduce all three core claims of Stevinson et al. from their toy classifier setting:PGD attacks against toy models generally agree with theoretically op