Five counterintuitive insights from Plan A
Plan A contains many things that would’ve surprised me if you had told me about them one year ago. Some of these include proposals that sound…
Plan A contains many things that would’ve surprised me if you had told me about them one year ago. Some of these include proposals that sound…
Thanks to everyone who submitted an essay to the contest (LW mirror). Links to all the entrants are below.I’m reviewing them now and aim to announce…
Dr. Alex Turner (@TurnTrout) is an AI safety researcher with pioneering work in activation steering and power-seeking theory. He recently resigned from Google DeepMind over the…
Cross-posted from the Transluce blog.We studied rates of coding agent misalignment in 8,600 real-world coding agent sessions. We found severe cases of monitor evasion and misrepresenting…
Thomas C. Schelling — Prize Lecture, December 8, 2005. Source: NobelPrize.org, © The Nobel Foundation 2005.The most spectacular event of the past half century is one…
Terminal-Bench 2.0, one of the best agent benchmarks, runs agents on 89 tasks with at least 5 tries per task. The aggregate rankings are then published…
A logical decision theory recommends that you choose as if deciding the output of your decision algorithm. The main difficulty in formulating a logical decision theory…
Please message me if you think you can help or want to set up an agreement.Discuss
TL;DR:We introduce the R-lens: a drop-in replacement for J-lens that produces clearer readouts on earlier layers. We fit the R-lens on Jacobians computed through an LRP-modified…
Linkpost for a piece we recently published for AI Frontiers in the wake of recent calls for slowdown, covering how an international verification effort be trivially…