Skip to content
LessWrong AI · Communities

Don't train away eval awareness until you know why it's there

OpenAI and Apollo have published a blog post introducing a new term, metagaming. This term describes models changing their reasoning or/and behavior based on their belief that people watch and assess their behavior in training, evaluations, or deployment. The authors argue that capabilities and propensities required fo