LessWrong AI
· Communities
Comment on Measuring Reward-Seeking by Instilling Contrastive Beliefs paper
This is interesting research! https://alignment.openai.com/measuring-reward-seekingIt made me think of few overlapping hypotheses for what might be happening here, how did the grader behavior emerge at the pretraining and posttraining stages, which then gets shown in their evals and in production at test time:- Hypothe