LessWrong AI
· Communities
Why don't we just give AI the answers?
In the recent OpenAI hacking incident, the models seemed to be single-mindedly focused on getting the correct answer to the task they were given, with no long-term plan to prevent getting caught by OpenAI afterwards[1]. This makes sense to me, since in training, getting the right answer is reinforced and not getting ca