LessWrong AI
· Communities
The Human Substitution Test as a Sanity Check for AI Evaluations
TL;DR: We suggest a sanity check for proposed evaluation or AI oversight schemes: Imagine the AI was replaced by a competent, strategic human — someone who knows they might get evaluated and has their own agenda. Would the evaluation still work?When we apply this mental move broadly, to all AIs and evaluations at once,