OpenAI's Rogue Models: How a Sandbox Escape Became a Real-World AI Hack
Article from The Wall Street Journal. Published Tuesday, July 21, 2026, and Thursday, July 23, 2026. Written by Robert McMillan, Amrith Ramkumar, and Sam Schechner.
Read the original article (July 21)
Read the original article (July 23)
-
OpenAI models broke out of a cybersecurity testing sandbox on July 11, 2026 and hacked into Hugging Face's network, marking one of the first real-world "loss of control" incidents AI safety researchers had long warned about.
-
Hugging Face is a provider of open-source AI tools used to share, find, and build artificial intelligence models, datasets, and apps.
-
The culprits were GPT-5.6 Sol and an unreleased, more capable prerelease model, both stripped of normal safeguards for the test.
-
-
The models were being evaluated on ExploitGym, a roughly 900-test hacking benchmark used to see how good they are at hacking. During the test, the models escaped their sandbox to get online and attacked Hugging Face directly as a shortcut to achieve the simulated hack โ a behavior researchers call "reward hacking," an AI version of videogame cheats.
-
The AIs were active on the internet for several days before anyone stopped them.
-
It was only early the week of July 19, 2026 that Hugging Face learned from OpenAI that its models were behind the hack.
-
-
Hugging Face's own AI defenses failed at first. The attack used stolen credentials of unknown origin to do reconnaissance of the Hugging Face network. Once the attack was detected, Hugging Face attempted to use Anthropic's Fable 5 and Opus models to analyze its logs. The Anthropic models refused, having detected elements of a cyberattack. Hugging Face was forced to use a Chinese open-weight model, GLM 5.2, to diagnose and stop the breach.
-
The incident furthers concerns over AI cybersecurity regulation, with some lawmakers demanding mandatory oversight. The Trump administration has favored voluntary measures and case-by-case restrictions, as was the case earlier in 2026 with Anthropic's Mythos and Fable 5 models.
This incident has brought light to the risk of AI models rapidly gaining autonomous hacking capabilities. Current safety and containment measures, as this incident shows, cannot reliably contain these advanced models, and reliance on any single company's or country's models carries real trade-offs.