LessWrong AI
· Communities
AI Mistake Seeding
I wonder if AI is being trained to make easy-to-correct mistakes so it can fix them later. That is, it ends up trained to correct its previous message's mistake, then make another mistake, so it can correct it again in the next message. From my understanding of RL, the human/AI judge has to rank several policy model re