LessWrong AI
· Communities
Transformers Resist Their Own Architecture
This is the first entry in a sequence of posts which compare a mathematical theory of attention against trained transformers.Links: [GitHub repository]: The code for these experiments. [Original paper]: Geshkovski et al., the theory this investigation is built on[YouTube walkthrough]: a video going over the original pa