Skip to content
LessWrong AI · Communities

Comment on Measuring Reward-Seeking by Instilling Contrastive Beliefs paper

This is interesting research! https://alignment.openai.com/measuring-reward-seekingIt made me think of few overlapping hypotheses for what might be happening here, how did the grader behavior emerge at the pretraining and posttraining stages, which then gets shown in their evals and in production at test time:- Hypothe