Skip to content
LessWrong AI · Communities

OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards

How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a