Skip to content
LessWrong AI · Communities

Transformers Resist Their Own Architecture

This is the first entry in a sequence of posts which compare a mathematical theory of attention against trained transformers.Links: [GitHub repository]: The code for these experiments. [Original paper]: Geshkovski et al., the theory this investigation is built on[YouTube walkthrough]: a video going over the original pa