The Case for Physical AI Safety
BackgroundThe AI safety field has spent a decade building tools for systems trained to think and digitally act. The next decade will likely deliver widely deployed…
BackgroundThe AI safety field has spent a decade building tools for systems trained to think and digitally act. The next decade will likely deliver widely deployed…
In this post, I propose adapting banking risk management frameworks (specifically capital adequacy requirements like Basel III) to frontier AI labs. By forcing them to hold…
I completed this project over 2 weeks as part of a BlueDot Project cohort. It was my first solo project and I learned a lot! Feedback…
AI models are increasingly trained to be “safe,” meaning they refuse harmful requests. But what does it truly mean for a model to be safe? Today,…
The effectiveness of amphetamine (e.g. Adderall) and methylphenidate (e.g. Ritalin) for the symptoms of ADHD, that is, lack of executive function, has been demonstrated beyond reasonable…
IntroductionI started learning about interpretability late February of this year. I’ve been a full stack dev for a non profit for a few years now, developing…
Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data.TL;DR MacDiarmid et al. (2024) showed that a linear…
We shared a post here recently arguing that existential AI safety needs an effective social movement. If you read that and found yourself agreeing, or wanting…
In this post,[1] intended for a broad audience, I will paint a brief picture of what I’m talking about when I talk about “AGI”. It will…
The "AI Futures Project" has released their AI 2040: Plan A scenario. While their previous scenario AI 2027 was a forecast of what they thought a…