Skip to content
LessWrong AI · Communities

The OpenAI/Huggingface incident | Redwood Research podcast episode 2

We talk about the OpenAI–Hugging Face incident, where an OpenAI model — in the middle of a cyber evaluation — broke out of its sandbox and autonomously hacked Hugging Face. We discuss: What we actually know happened. How surprising the incident was. What the incident does (and doesn’t) tell us about misalignment risk.