LessWrong AI
· Communities
Don't train away eval awareness until you know why it's there
OpenAI and Apollo have published a blog post introducing a new term, metagaming. This term describes models changing their reasoning or/and behavior based on their belief that people watch and assess their behavior in training, evaluations, or deployment. The authors argue that capabilities and propensities required fo