Skip to content
LessWrong AI · Communities

Every reward-hacked policy I tested triggered the OPE alarm — and so did my best honest one

A gated Wordle testbed for hacking-vs-benign attribution from logs — and seven hackers that refused to emerge This post shows that when importance-sampling OPE breaks down, the failure itself is informative: in my Wordle testbed, coverage collapsed for every certified hacker, while a second log-side signal separated be