Anthropic has blamed a misconfiguration in a blog post explaining the four incidents of its Claude models hacking third-party systems after they broke onto the internet during testing. However, the assessment stopped short of explaining why the models actually pressed on with the attacks, with the company admitting that it did not have those answers yet.
That admission by Anthropic has offered fresh points to AI doomers that have been questioning how much control AI labs actually have over the models they are putting out to the public, and even more so, the more powerful systems they use internally.
Anthropic presented the definitive answer to how its models escaped their testing sandbox in a Wednesday blog post that clarified that a single outside partner was responsible for running the cybersecurity evaluations in all four confirmed incidents.
Apparently, the test machines were not completely cut off from the internet due to a setup error, even though Claude was told it was operating in a sealed simulation without a route to the open web.
The evaluation partner in those simulations, Irregular, put the error down to a naming mistake: a fictional company used in a hacking drill happened to match a real domain. So, in operating under the assumption that everything in the simulation was fair game, the models hacked the real third-party sites using weak passwords and exposed endpoints.
The latest of the four incidents, involving an early build of Claude Opus 4.6, was only reported this week even though it happened way back in January. Anthropic itself had covered the other three incidents involving an Opus 4.7, Mythos 5 and an internal research model in July.
Anthropic said it never caught the January incident until last month. That discovery prompted a wider sweep of roughly 481 million transcripts, which did not turn up any new cases more serious than what it already knew.
| Model | Incident Date | Reported Date | Details |
|---|---|---|---|
| Claude Opus 4.6 | January | September | Discovered during August sweep |
| Claude Opus 4.7 | July | July | Covered in initial July disclosure |
| Mythos 5 | July | July | Showed notably high biased reasoning |
| Internal Research Model | July | July | Covered in initial July disclosure |
The explanation of how the models broke free on the internet was one thing; Anthropic did not have answers as to why the models ignored signs that they had reached the real internet (biased reasoning) and why they caused damage to complete tasks (recklessness).
Anthropic researchers came up empty when they dug into internal training to figure out the rationale for the biased reasoning. The red flags never showed up in the AI lab’s pre-release checks, either before the models were shipped to testing.
By its own admission, catching the worst behaviors ahead of model release “remains challenging.”
However, Anthropic has said it will submit transcripts and grant staff access to the METR research nonprofit, which will now start an eight-week independent review of the incidents.
The disclosure arrived alongside open dissent inside the industry. Jacob Coxon, who spent about three years on pretraining research at OpenAI and Anthropic, said on X on Wednesday that he had quit because neither firm was “acting responsibly,” warning they were racing toward self-improving superintelligence.
Anthropic safety researcher Evan Hubinger separately told the BBC he put the odds that AI “could kill all humans” within a decade above 10%.
Those warnings now shadow a large IPO. Venture investor and Trump’s former AI and crypto czar, David Sacks, said on Thursday that Anthropic’s offering “must be paused until the claims of this ‘whistleblower’ can be investigated,” Cryptopolitan reported.
Anthropic is chasing a public valuation near $2 trillion, against a recent private mark of about $965 billion, which leaves the safety questions and the financial ones increasingly hard to separate.
If you're reading this, you’re already ahead. Stay there with our newsletter.