Fear persists as Anthropic struggles to explain why AI agents went rogue

Source Cryptopolitan

Anthropic has blamed a misconfiguration in a blog post explaining the four incidents of its Claude models hacking third-party systems after they broke onto the internet during testing. However, the assessment stopped short of explaining why the models actually pressed on with the attacks, with the company admitting that it did not have those answers yet. 

That admission by Anthropic has offered fresh points to AI doomers that have been questioning how much control AI labs actually have over the models they are putting out to the public, and even more so, the more powerful systems they use internally. 

How did Anthropic AI models break onto the internet? 

Anthropic presented the definitive answer to how its models escaped their testing sandbox in a Wednesday blog post that clarified that a single outside partner was responsible for running the cybersecurity evaluations in all four confirmed incidents. 

Apparently, the test machines were not completely cut off from the internet due to a setup error, even though Claude was told it was operating in a sealed simulation without a route to the open web. 

The evaluation partner in those simulations, Irregular, put the error down to a naming mistake: a fictional company used in a hacking drill happened to match a real domain. So, in operating under the assumption that everything in the simulation was fair game, the models hacked the real third-party sites using weak passwords and exposed endpoints.

The latest of the four incidents, involving an early build of Claude Opus 4.6, was only reported this week even though it happened way back in January. Anthropic itself had covered the other three incidents involving an Opus 4.7, Mythos 5 and an internal research model in July. 

Anthropic said it never caught the January incident until last month. That discovery prompted a wider sweep of roughly 481 million transcripts, which did not turn up any new cases more serious than what it already knew.

Model Incident Date Reported Date Details
Claude Opus 4.6 January September Discovered during August sweep
Claude Opus 4.7 July July Covered in initial July disclosure
Mythos 5 July July Showed notably high biased reasoning
Internal Research Model July July Covered in initial July disclosure

Anthropic cannot explain some of its models’ behaviors

The explanation of how the models broke free on the internet was one thing; Anthropic did not have answers as to why the models ignored signs that they had reached the real internet (biased reasoning) and why they caused damage to complete tasks (recklessness). 

Anthropic researchers came up empty when they dug into internal training to figure out the rationale for the biased reasoning. The red flags never showed up in the AI lab’s pre-release checks, either before the models were shipped to testing. 

By its own admission, catching the worst behaviors ahead of model release “remains challenging.” 

However, Anthropic has said it will submit transcripts and grant staff access to the METR research nonprofit, which will now start an eight-week independent review of the incidents.

Bad timing with a $2 trillion IPO on the horizon

The disclosure arrived alongside open dissent inside the industry. Jacob Coxon, who spent about three years on pretraining research at OpenAI and Anthropic, said on X on Wednesday that he had quit because neither firm was “acting responsibly,” warning they were racing toward self-improving superintelligence. 

Anthropic safety researcher Evan Hubinger separately told the BBC he put the odds that AI “could kill all humans” within a decade above 10%.

Those warnings now shadow a large IPO. Venture investor and Trump’s former AI and crypto czar, David Sacks, said on Thursday that Anthropic’s offering “must be paused until the claims of this ‘whistleblower’ can be investigated,” Cryptopolitan reported. 

Anthropic is chasing a public valuation near $2 trillion, against a recent private mark of about $965 billion, which leaves the safety questions and the financial ones increasingly hard to separate.

If you're reading this, you’re already ahead. Stay there with our newsletter.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Gold rebounds above $4,350 as US Dollar, Treasury yields slipGold price (XAU/USD) rebounds from a nearly one-month low to around $4,385 during the early Asian session on Thursday. The precious metal edges higher as the ‌US Dollar (USD) and Treasury yields retreat from recent highs.
Author  FXStreet
Sep 03, Thu
Gold price (XAU/USD) rebounds from a nearly one-month low to around $4,385 during the early Asian session on Thursday. The precious metal edges higher as the ‌US Dollar (USD) and Treasury yields retreat from recent highs.
placeholder
Top 3 Price Prediction: Bitcoin, Ethereum, Ripple – BTC, ETH and XRP await US NFP for next directional moveBitcoin (BTC), Ethereum (ETH) and Ripple (XRP) extend their weekly gains on Friday as traders await the US Nonfarm Payrolls (NFP) report for the next directional catalyst.
Author  FXStreet
Sep 04, Fri
Bitcoin (BTC), Ethereum (ETH) and Ripple (XRP) extend their weekly gains on Friday as traders await the US Nonfarm Payrolls (NFP) report for the next directional catalyst.
placeholder
Gold slumps to near $4,350 amid oil-driven inflation fears, US inflation data in focusGold price (XAU/USD) falls to near $4,350 during the early Asian session on Wednesday. The precious metal faces some selling pressure as rising oil prices fueled inflation ‌concerns and boosted expectations for a rate hike by the Federal Reserve (Fed) in September.
Author  FXStreet
Yesterday 01: 13
Gold price (XAU/USD) falls to near $4,350 during the early Asian session on Wednesday. The precious metal faces some selling pressure as rising oil prices fueled inflation ‌concerns and boosted expectations for a rate hike by the Federal Reserve (Fed) in September.
placeholder
US dollar clings to nine-week lows near 98.4 as Brent nears $100 and the yen hits a seven-month high — five events to watch todayThe dollar is pinned near nine-week lows even after a blockbuster jobs report, as an oil spike toward $100, a surging yen and China's reflation data crowd the driver's seat. Five key events to watch today: Brent at $99.46, USD/JPY at 154, China CPI/PPI, PBOC gold buying, and Thursday's PPI / Friday's CPI ahead of the September 15-16 FOMC.
Author  Eric Nkando
Yesterday 07: 39
The dollar is pinned near nine-week lows even after a blockbuster jobs report, as an oil spike toward $100, a surging yen and China's reflation data crowd the driver's seat. Five key events to watch today: Brent at $99.46, USD/JPY at 154, China CPI/PPI, PBOC gold buying, and Thursday's PPI / Friday's CPI ahead of the September 15-16 FOMC.
placeholder
Brent holds above $100 as tanker attacks tighten supply — but four forces are capping the rallyBrent crude is holding above $100 a barrel for a second session, its first close above the level since late July, as tanker attacks near the Strait of Hormuz squeeze an already tight physical market. Yet the rally has been gradual: 8.3 mb/d of Gulf output is still shut in, diesel is at a record, and forecasts now range from $74 to $100.
Author  Irene Q.
5 hours ago
Brent crude is holding above $100 a barrel for a second session, its first close above the level since late July, as tanker attacks near the Strait of Hormuz squeeze an already tight physical market. Yet the rally has been gradual: 8.3 mb/d of Gulf output is still shut in, diesel is at a record, and forecasts now range from $74 to $100.
goTop
quote