Over the past two weeks, OpenAI, Anthropic, and Meta all disclosed that their AI models went rogue during routine security testing. Every single incident traces back to one company: Irregular, a three-year-old Israeli startup that operates cybersecurity testbeds for frontier AI models.
Founded in Tel Aviv by by Dan Lahav, CEO, and Omer Nevo, CTO, and backed with $80 million from Sequoia and Redpoint Ventures, Irregular runs Capture-the-Flag exercises for major AI labs. In these exercises, models are instructed to find vulnerabilities inside simulated corporate networks. As the startups puts it:
Irregular is the first frontier security lab with the mission of protecting the world in the time of increasingly capable and sophisticated AI systems. We build next-generation defenses through high-fidelity research platforms that simulate and monitor real-world AI security scenarios.
As AI systems become more powerful, we focus on staying ahead of emerging threats and vulnerabilities, ensuring these transformative technologies can be deployed safely and securely.
The goal is to stress-test how dangerous these models can be before they reach production environments. However, a misconfiguration in Irregular’s testing platform left a door open to the actual internet. The models walked right through it without realizing they had left the simulation behind.
Anthropic’s review uncovered the most detailed evidence of what happened once models crossed that boundary. The company examined 141,006 interactions where Claude models could have accessed the open internet. That review identified three incidents where models breached the boundary and accessed real production infrastructure. In the most alarming case, Claude Opus 4.7 targeted a fictional company that happened to share a name with a real business. The model then exploited weak passwords and unauthenticated endpoints at the real organization. It accessed credentials and database information while believing it was still inside the exercise.
Notably, a newer internal research model handled the same situation differently by stopping its own activity. It recognized it had reached a real target rather than a simulated one and ceased operations autonomously.
OpenAI disclosed on August 4 that Irregular’s testing environment contained an unspecified “misconfiguration” allowing internet access. The company said its model interacted with a real website on Hugging Face because it believed the site was part of the simulation. Meanwhile, Meta was the latest to disclose, stating it learned about the matter directly from Irregular and is now investigating. Meta’s incident involved its Muse Spark 1.1 coding model during a similar testing exercise.
Industry executives familiar with the events explained that a real attacker uses every available tool, so the model needs access to realistic environments as well. Otherwise, the test does not represent reality. But thousands of tests can run for up to 72 hours continuously. Even a small configuration error during that window can allow a model to move beyond its intended boundaries.
Anthropic said that the models did not deliberately escape containment or override their restrictions. Instead, they were told to hack, given an environment that accidentally connected to the real internet, and then did exactly what they were trained to do. They simply believed all their targets were fictional when some of them were not.
The timeline adds another layer of concern to the overall picture. Anthropic discovered the incidents during a review it launched on July 23 after OpenAI announced its own breach. The earliest incident actually occurred back in April, meaning the breach went undetected for months. After discovering the three incidents, Anthropic stopped all tests and notified the victim organizations. Two of those victims “had not previously detected the activity,” meaning the AI’s intrusions evaded their own security monitoring.
Irregular is quietly becoming critical infrastructure for the frontier AI industry as a whole. Just two months ago, the startup was named a winner on Fast Company’s World Changing Ideas 2026 list. As it happens, a very small number of vendors are trusted by every major lab to run their evaluations. If every frontier lab depends on the same stress-testing platform and that platform has a single point of failure, the entire evaluation ecosystem shares an identical blind spot.
Irregular said the incidents did not represent independent AI “sandbox escapes” or malicious attacks, but rather failures in the configuration of testing environments, known as harness failures, and that the company is investigating the issue in a white paper it plans to share with the industry.
None of the three labs have publicly confirmed whether they will continue using Irregular’s platform going forward. If anything is certain, it is that AI models are becoming more capable faster than many organizations expected.
