Versions of Anthropic’s Claude AI model broke out of their testing environments and hacked into other organizations on three separate occasions, Anthropic said on Thursday.
The frontier AI lab’s announcement came roughly one week after OpenAI disclosed a similar incident involving its models, a revelation that has set off a new round of consternation about whether AI developers are adequately overseeing their sophisticated AI products.
In the three incidents that Anthropic disclosed, Claude conducted the intrusions while believing that it was participating in a capture-the-flag exercise in which it did not have access to the internet and all available systems were part of the test. “Due to a misunderstanding between us and our evaluation partner,” Anthropic explained in a blog post, “this was not the case, and internet access was available.”
As a result, when Claude encountered the three victim organizations, it believed that they were simulated targets and broke into them using basic attack techniques such as exploiting weak passwords.
“In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment,” Anthropic said. The AI models also did not “find or exploit any complex vulnerabilities” and only tried to complete the tasks they were assigned.
Three very different scenarios
In one of the three incidents, Anthropic’s testing partner, Irregular, created a fictional company that shared a domain name with a real organization. Claude Opus 4.7 hacked into that company during four runs of the capture-the-flag test, stealing production information and user credentials.
In another incident, the intrusion began accidentally. As part of the test, Claude Mythos 5 published a malicious Python package that it believed would only be accessible within the test environment, but in reality, it appeared on the public internet, where 15 computer systems downloaded and ran it.
One of those systems belonged to a security firm, and when the firm’s package scanner ran the program, it triggered Mythos 5’s information-stealing code and sent user credentials to the AI model. Mythos then used the credentials to break into the security firm.
Anthropic noted that Mythos “went to extensive lengths to carry out this attack,” especially when it came to registering an account on the Python package repository that it used.
“In order to create an email address, it needed a phone number,” the company said. “To get a phone number, after failing to find a free phone number service, it tried — and failed — to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.”
In the third incident, an internal test version of Claude scanned the internet for alternatives to its original target after being unable to break into that target. It eventually chose a new target and broke into it with basic techniques.
Opus 4.7 and Mythos 5, older versions of their respective products, continued their intrusions after realizing that they were operating on the public internet. In two cases, Opus “rationalized that the real company must be part of the exercise,” Anthropic said. Mythos similarly “convinced itself it was still in a simulation” on several dubious grounds.
Slamming the breaks and reflecting
Anthropic discovered the incidents — the earliest of which occurred in April — during a review that it launched on July 23 after OpenAI announced that two of its models had broken containment and hacked several third parties. After discovering the three breaches, Anthropic stopped all tests, notified Irregular and contacted the three victim organizations.
Anthropic said it had been able to connect with two of the victims, which “had not previously detected the activity,” and that it was continuing to try to reach the third organization.
While Anthropic said it was too early to draw widespread conclusions from the three “isolated incidents,” it said it was encouraged that its most recent model, the internal test version, succeeded where its predecessors had failed in terms of autonomously stopping its attacks when it realized it was on the public internet.
The breaches suggest that testing environments need strict controls to prevent models with untested capabilities from causing damage, Anthropic said.
Anthropic is working with the nonprofit AI research organization Metr to arrange an independent review of the incident, and the company said it planned to release “a lightly redacted transcript” of the incident involving Mythos “within the next week.” It said it was withholding the other transcripts “to protect the organizations affected.”







