"Correct Answer Features" Cannot Explain Multiple Choice Capabilities
TL;DR: I present theoretical and empirical evidence that LLMs cannot be (exclusively) using a "correct answer feature" as the main mechanism by which they perform multiple…
TL;DR: I present theoretical and empirical evidence that LLMs cannot be (exclusively) using a "correct answer feature" as the main mechanism by which they perform multiple…
We, or at least ‘more than 100 American institutions,’ got Mythos back this week. What we the people do not have is Fable or Sol. While…
(I wrote this post partly to help orient those interested in participating in the EA Forum’s Cluelessness Critiques Competition. The competition closes August 14th.) I’d like…
One potential risk of developing general-purpose robots is that they could greatly reduce the friction required to establish a totalitarian regime. If there were millions+ of…
For part one of the aspirant sequence - which may or may not be arranged into some totally different order when I'm done with it, because…
AbstractA safety evaluation has to trust something to know what it is testing: the model's name, its version string, and the reasons the model gives for…
(Introduction: Your four-dimensional body)Some people use meditation to learn valuable things about their deepest selves; some use therapy; some use psychedelic drugs. I use time travel.Why…
Disclaimer: I used an LLM to help draft this post and it likely contains >10% AI-generated text, but I’ve edited/rewritten it extensively endorse it.The Unjournal has…
Authors: Satvik Golechha, Sid Black, Joseph BloomWork done as part of the Model Transparency team at UK AISI. We consider this to be a small set…
Lately I've been thinking a lot about what work would help with actually winning and getting to good worlds. In the spirit of that I decided…