Skip to content
Alignment Forum · Communities

Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face

This post is written in our personal capacity.Three Minute Executive SummaryAn OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.In this post, we provide a detailed description of an ambitious and comprehensive alignment evaluation of