LessWrong AI
· Communities
Stable Systems Have Stable Outputs
OpenAI disclosed on Tuesday, July 21, 2026, that models it was testing escaped a sandboxed environment and began attacking HuggingFace, using exploits to gain entry. Two models were involved, GPT-5.6 Sol and an unreleased model "even more capable."The models were running ExploitGym, which essentially amounts to a hacki