Skip to content
LessWrong AI · Communities

Why study proto-training gaming as an adversarial alignment failure mode?

This is a dual post that lays out our current research project where we compare different pre-RL alignment methods and their ability to prevent models from ‘proto-training gaming,’ which we predict is selected for over the course of RL post-training. In the previous post, we enumerated possible pre-RL alignment interve