LessWrong AI
· Communities
Making Credible Deals With AI
If we end up with a weakly superhuman scheming AI, trading with it may reduce takeover risk: offer it things it values (donate to causes that further its goals) in exchange for useful behavior (revealing misalignment, better alignment auditing techniques). Others have argued the case[1].One key bottleneck is credibilit