On Democratizing ASI to Preserve Civil Liberties
I continue to believe we should pause frontier AI development. Any discussion of alternative strategies should be thought of as planning for contingencies. A unifying driver…
I continue to believe we should pause frontier AI development. Any discussion of alternative strategies should be thought of as planning for contingencies. A unifying driver…
Code and data available at github.com/KieronKretschmar/latent-awarenessTL;DRWe take two eval-gaming model organisms (Hua et al.'s (2025) organism and RogueQwen) and apply direct preference optimization (DPO) to their…
One potential risk of developing general-purpose robots is that they could greatly reduce the friction required to establish a totalitarian regime. If robots became physically capable…
This is the advice I wish I had when I started trying to become an AI safety research engineer.The LandscapeStart by working out which issues you…
For... idk, at least a few months? I think a briefly useful job is "Guy who basically enters Claude Code prompts for you, but, manages ironing…
I try to think about topics like desire, causation, and evidence and often find myself in puddles of confusion. I’ve recently wondered how, when someone can…
Epistemic status: I consider the following future quite plausible in the next few years (~35% chance that something vaguely like this occurs), perhaps as soon as…
What we did, in a sentence: we built an index that attempts to score countries by their current vulnerability to AI-amplified democratic backsliding. It’s an early…
If you take the weight difference between a reasoning model and its non-reasoning instruct counterpart, and then apply more of that difference to the reasoning model,…
C: Babe, whatever happens, I really appreciate you doing this for me.J: Okay. I still don’t think it’s a good idea.C: Look, it’s a one-time thing.…