Watch a chess transformer think
I was poking around at a chessformer which mimics human play and made this fun companion app to visualize a lot of the internals of the…
I was poking around at a chessformer which mimics human play and made this fun companion app to visualize a lot of the internals of the…
Related: Proposal for making credible commitments to AIs Making deals with early schemers Establishing credibility is the baseline for trust; trust in turn enables (richer) bargaining.…
IntroductionOver the past few years, AI tools have become useful for conducting technical AI research. In the early ChatGPT era (~2023–2024), chat assistants were maybe useful…
I previously have written back in March 2022 about how I use Twitter, and back in April 2023 about Twitter and its then-new algorithms, which have…
This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models. It gives the intro, main results table, and example prompts from the…
This is a continuation of the post Your Brain Has an Attack Surface. If you haven’t read it, here is the short version: there is a…
About a year ago, I began transitioning from software engineering to AI safety research. I was drawn into this by a question that arose while building…
Last year, I wrote an essay "The hard problem of qualia in the age of AI", https://zenodo.org/records/20549564 (11 pages PDF). I want to create a linkpost…
LLMs are commonly assumed to use superposition to represent more features than they have dimensions. The evidence for this is mostly indirect — chiefly the success…
We propose synthetic scalable oversight, a technique for studying scalable oversight by creating graphical abstractions of real-world problems and training tiny models inside these synthetic environments…