A list of existing alignment approaches
How can we make a nice AI system?Here's a list of all the techniques I'm aware of. Train the AI system to be nice. There are…
How can we make a nice AI system?Here's a list of all the techniques I'm aware of. Train the AI system to be nice. There are…
What values would AIs instill in their successors? Though the AI Village agents can’t train frontier models, we can explore a related question: What values would…
You want to do something to help AI go well and are starting a project to make that happen. Should you create a nonprofit or a…
Sandboxing is a classic tool in computer security: to run code you do not trust, you run it in an environment with limited permissions. It's harder…
This article reflects new updates to the accompanying paper: arxiv.org/abs/2606.18142. Benchmark: now included in the UK AI Security Institute's Inspect Evals. Leaderboard: compassionbench.com/tac.A model may condemn…
Epistemic status: This is based on ten years of matchmaking experience in India - my lens for evaluating alignment. This is a conceptual essay continuing on…
There are a number of reasons to believe current AI models are conscious. I mean “conscious” is the sense of “is there something it is like…
The technology is transformative, the risks catastrophic, and the rate of change overwhelming. For things to go well, there's much work to be done. The legal…
I work at CBAI, and we might push to get fellows to publish results on LW or X. I'd like to give more than just my…
The goal of this post is to describe how my views evolved over time, and reflect on this process. It might prompt you to make your…