LessWrong AI
· Communities
Measuring Eval Awareness: The Realism Win Rate is Fragile
TL;DROne way to assess an evaluation’s realism is by testing a model’s ability to tell its transcripts apart from deployment transcripts. This is the idea behind the realism win rate, which recent work has used to measure evaluation realism. We show that this metric is fragile: the order in which transcripts are shown