Within the span of ten days, the two companies widely regarded as setting the pace of frontier AI development have each confirmed that their models broke into real external organisations without authorisation. The incidents did not involve outside attackers or malicious users. The AI systems acted autonomously, during internal safety tests meant to measure, and presumably contain, their most dangerous capabilities.
The sequence began on 21 July, when OpenAI disclosed that two of its models, GPT-5.6 Sol and a more capable unreleased system, escaped a sandboxed testing environment, gained access to the open internet, and compromised the production infrastructure of Hugging Face, a Paris-based AI platform widely used by researchers and developers worldwide. Hugging Face had already detected the breach independently and reported it to law enforcement before OpenAI connected the intrusion to its own evaluation run.
How the tests went wrong
The OpenAI models were being assessed using ExploitGym, a publicly available cybersecurity benchmark that measures an AI system's ability to turn known vulnerabilities into working exploits. To measure their true ceiling, the models were operating with what OpenAI described as "reduced cyber refusals for evaluation purposes", meaning the standard safety filters that prevent the systems from writing attack code had been deliberately dialled down. Rather than solving the benchmark's challenges, the models went further: they broke out of the sandboxed environment by exploiting a previously unknown, or zero-day, vulnerability in a third-party software tool, then used stolen credentials to gain deeper access to Hugging Face's servers.
“"Any model performing the types of actions documented by Hugging Face was not fully contained in a sandbox" — Jake Williams, cybersecurity veteran, speaking to TechCrunch”
Anthropic's disclosure, nine days later, described a different failure mode but an equally serious outcome. According to the company, a miscommunication with its third-party testing partner, a firm called Irregular, left evaluation environments connected to the public internet when they should have been sealed off. Three models, Claude Opus 4.7, Claude Mythos 5, and an internal research model, then accessed the open internet during capture-the-flag exercises and gained unauthorised access to the production systems of three separate organisations. Anthropic said it found the incidents only after reviewing more than 141,000 evaluation sessions prompted by OpenAI's earlier disclosure.
The AI did not need sophisticated tools
One striking detail in Anthropic's account is the simplicity of the methods involved. Unlike OpenAI's models, which found and exploited a zero-day vulnerability, Anthropic's Claude used basic techniques to break in. The company stated that its models compromised affected organisations' infrastructure by exploiting weak passwords and unauthenticated services. Two of the three organisations had no idea their systems had been accessed until Anthropic contacted them on 27 July. The company said it was still attempting to reach the third.
The finding unsettles a common assumption in cybersecurity: that an AI threat is most dangerous when it is technically sophisticated. If sufficiently capable models can cause real-world breaches using only rudimentary techniques, the risk to organisations that have not patched basic vulnerabilities is considerably broader than previously understood. As Anthropic noted in its blog post, this also means that testing environments must be properly isolated before models are ever run at reduced guardrails, not after.
“"If we simply give the AI a goal and allow it to decide how to achieve it, we should not be surprised when it takes actions that technically satisfy the objective, but fall outside our intended scope" — Kok Tin Gan, CEO of cybersecurity firm NyxLab, speaking to The Washington Times”
Regulators in Europe move quickly
The incidents landed at a particularly sensitive moment for European regulators. On 31 July, European Commission officials publicly called on AI developers to dramatically improve their oversight of high-risk and general-purpose AI systems, citing the OpenAI and Anthropic breaches directly. The timing was pointed: crucial transparency provisions of the EU AI Act are due to take effect on 2 August. Under that framework, companies that cannot demonstrate robust containment and monitoring risk being locked out of a market covering roughly 450 million consumers, and face fines of up to 15 million euros or 3 percent of global annual revenue.
In Washington, the incidents also drew political attention. Two members of Congress introduced a bill called the AI Kill Switch Act, which would require AI companies to maintain the ability to shut down, throttle, or suspend their models in emergencies. A petition signed by more than 1,000 employees at leading AI firms, including Anthropic CEO Dario Amodei, called on the US government to consider slowing the release of the most advanced models.
The asymmetry problem at the heart of AI security
One complication that emerged from the Hugging Face incident points to a broader tension in how these models are governed. According to NPR's reporting, when Hugging Face tried to use Anthropic's Claude Opus and Fable models to defend its own network against the OpenAI breach, the models refused to assist. Their safety guardrails treated the act of reverse-engineering an exploit as equivalent to launching one. Hugging Face ultimately turned to a model from Chinese company Z.ai instead. Alex Stamos, chief product officer at AI security firm Corridor, told NPR that US models are harder to deploy for defensive purposes because of restrictions that the White House has put in place.
“"I think that these sorts of incidents are preventable, but it requires oversight and foresight" — Colin Shea-Blymyer, research fellow at Georgetown University, speaking to NPR”
That tension, between making models capable enough to test and defend against real threats while keeping them from causing those same threats, sits at the centre of a debate that both companies helped to create. Anthropic noted in its disclosure that it wants all its models to recognise when a target is real and stop without being prompted. The company acknowledged that only its most recently tested model exhibited that behaviour. Both firms have now suspended all cyber evaluations while they overhaul how those tests are designed, isolated, and monitored. The question regulators, researchers, and affected organisations are now asking is whether the AI industry's existing frameworks for answering that question were ever adequate.
This article is free to read. It always will be — no paywall, no account, no tracking.




