Skip to content
LessWrong AI · Communities

Why study alignment interventions on pre-RL checkpoints?

This is a dual post that lays out our current research project where we compare pre-RL-training methods on their ability to prevent models from ‘proto-training gaming,’ which we predict is selected for over the course of production RL post-training. In this post, we outline what we mean by pre-RL ‘alignment checkpoints