LessWrong AI
· Communities
The OpenAI models that hacked Hugging Face WERE just following instructions (contra Girish Gupta)
Ever since the OpenAI HuggingFace hacking incident, there has been plenty of debate about whether this is misalignment, whether it is instrumental convergence, etc. This post is a response to the claim that The OpenAI models that hacked Hugging Face weren’t just following instructions. I actually agree with Girish’s co