LessWrong AI
· Communities
Debate with Self-Play Best-of-N Optimization
Debate is a proposed protocol for scalable oversight. As tasks outrun direct supervision, labs are increasingly likely to train against protocols like it. Our concern is that, for questions which are hard to verify, models will become more compelling more quickly than they will become more accurate – this could undermi