The Human Substitution Test as a Sanity Check for AI Evaluations
TL;DR: We suggest a sanity check for proposed evaluation or AI oversight schemes: Imagine the AI was replaced by a competent, strategic human — someone who…
TL;DR: We suggest a sanity check for proposed evaluation or AI oversight schemes: Imagine the AI was replaced by a competent, strategic human — someone who…
TL; DR: I find it hard to believe that сap-and-trade from AI-2040 would be effective in a context involving prohibitions. This will likely create an unsustainable…
A group is worried about an approaching fire spreading rapidly through their city. They manage to halt the fire outside the city gates. Meanwhile they build…
This is part 2 of the weekly, broadly covering speculation, rhetoric and policy, along with alignment research. This does not cover the release of GPT-5.6-Sol. As…
See 1 Jan 2026, 1 Jan 2025 and the July 2025 update. This continues my habit of documenting my beliefs and feelings as we transition to…
Part 1 of a series on process alignment, a different way of thinking about agents at every scale.IntroductionThere are quite a lot of disagreements in the…
I firmly believe that value generalisation[1]is the key to AI Alignment. That, indeed, it is necessary and almost sufficient for alignment.But I won't be arguing that…
Epistemic status: my first solo post here, so critiques of both substance and form are welcome. Also, note that, while I wrote this and all of…
Reading and editing the visual workspace of a vision-language model.Vision-language models hallucinate: ask one whether some object is in a picture and it will happily say…
This post serves to argue that backdooring evaluations are prone to failures stemming from triggers never reaching models.In backdooring literature, there is a common workflow. Outputs…