Skip to content
LessWrong AI · Communities

A list of existing alignment approaches

How can we make a nice AI system?Here's a list of all the techniques I'm aware of. Train the AI system to be nice. There are a variety of things we can vary in how we train the AI:Train using model internals OR using outputs.The central internals-based things I’m imagining involve using the internals as a reward signal