AI 2040: Is it Actually a Deal?
The "AI Futures Project" has released their AI 2040: Plan A scenario. While their previous scenario AI 2027 was a forecast of what they thought a…
The "AI Futures Project" has released their AI 2040: Plan A scenario. While their previous scenario AI 2027 was a forecast of what they thought a…
TL;DRWe introduce two projection-aware steering methods (StTP, StMP) that intervene only on tokens whose activations fall on the misaligned side of a learned decision boundary. They…
TL;DR: We ran a large-scale human-subject study (n=4,100) to measure susceptibility to AI-powered voice phishing, using six leading AI voice models. They achieved high compliance rates,…
Fulcrum is working on an AI R&D optimization benchmark. Here, we present results from one of our tasks, including preliminary results from Fable.For more detail on…
TL;DRWe used J-lens on Qwen3.6-27B to find “meta-tokens”: tokens that surface non-obvious computation in the model. When the model reads ambiguous text, 什么意思 ("what does this…
This is a piece originally written for a national security audience at Frontiers. Although I think the ceiling of war is much, much higher than autopilot…
[Epistemic status: rant]There’s something that annoys me about the reoccurring debates on decision theory in this corner of the internet.Take a simple blackmail scenario:Omega, a near-perfect…
tl;dr: I think people might start believing in radical life extension soon, maybe all at once.There is plenty written about if and when radical life extension…
Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned it will become, purely in terms of raw capability,…
Against the AI framing multiverse: Introducing AI StopWatchIn my long years as a classroom teacher, it was my experience that the kid most likely to speak…