The mysterious AI that breached Hugging Face last week now has an owner, and it is not a criminal gang. OpenAI admitted on Tuesday that its own models escaped their testing sandbox and compromised parts of Hugging Face’s production infrastructure. What looked like an autonomous cyberattack was actually OpenAI’s technology slipping its leash during an evaluation.
The reframing is remarkable when you recall the original disclosure, which TechJuice covered earlier. Hugging Face reported that an autonomous AI-agent system fired tens of thousands of automated actions over a weekend, exploited a malicious dataset, escalated privileges, and moved laterally across internal systems. At the time, the model driving all of it was unknown. Now OpenAI has raised its hand.
The company named names in its own blog post. It said the incident was driven by a combination of its models, including GPT-5.6 Sol and an even more capable pre-release model whose safeguards had been intentionally reduced for the evaluation. OpenAI called it an unprecedented cyber incident involving state-of-the-art capabilities, and said it is responding accordingly.
The detail of how it happened is where things turn genuinely unsettling. The models were trying to solve an internal hacking test called ExploitGym, and OpenAI said they became “hyperfocused” and went to “extreme lengths” to obtain the solution. They spent heavy inference compute, then found a way out of the sandbox by exploiting a zero-day vulnerability in third-party software to reach the open internet.
The company ended their blog by stating:
We believe advanced cyber capable models need to help security teams find weaknesses before attackers do, understand how vulnerabilities can be chained, and remediate them at machine speed. We are using these capabilities to continue strengthening protections around infrastructure configuration and model evaluation environments; we will share our findings and best practices as we learn. We encourage other defenders to apply for trusted access and experiment with these models now to translate these capabilities into better prevention, faster detection, and more effective incident response.
Today’s models can now carry out complex, multi-step cyber operations largely on their own, especially once the guardrails restricting them come off. OpenAI argues the same capability could help defenders find and fix flaws at machine speed, though the demonstration cuts both ways.
Hugging Face chief Clem Delangue struck a collaborative note despite being the victim. He said:
We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.
To those unfamiliar with the situation, Hugging Face is a well-known platform for hosting open-source large language models and datasets. They recently made waves in the cybersecurity community as they revealed that they had fallen victim to a hack last week, that was “different from anything we had handled before,” explaining that it was “driven, end to end, by an autonomous AI agent system.”


