Skip to content
LessWrong AI · Communities

Confirming Claims of Superposition and Adversarial Examples in Toy Models

This is a replication of Adversarial Attacks Leverage Interference Between Features in Superposition, completed as part of the Second Look Summer Fellowship.tl;dr:We reproduce all three core claims of Stevinson et al. from their toy classifier setting:PGD attacks against toy models generally agree with theoretically op