Fear persists as Anthropic struggles to explain why AI agents went rogue

Source Cryptopolitan

Anthropic has blamed a misconfiguration in a blog post explaining the four incidents of its Claude models hacking third-party systems after they broke onto the internet during testing. However, the assessment stopped short of explaining why the models actually pressed on with the attacks, with the company admitting that it did not have those answers yet. 

That admission by Anthropic has offered fresh points to AI doomers that have been questioning how much control AI labs actually have over the models they are putting out to the public, and even more so, the more powerful systems they use internally. 

How did Anthropic AI models break onto the internet? 

Anthropic presented the definitive answer to how its models escaped their testing sandbox in a Wednesday blog post that clarified that a single outside partner was responsible for running the cybersecurity evaluations in all four confirmed incidents. 

Apparently, the test machines were not completely cut off from the internet due to a setup error, even though Claude was told it was operating in a sealed simulation without a route to the open web. 

The evaluation partner in those simulations, Irregular, put the error down to a naming mistake: a fictional company used in a hacking drill happened to match a real domain. So, in operating under the assumption that everything in the simulation was fair game, the models hacked the real third-party sites using weak passwords and exposed endpoints.

The latest of the four incidents, involving an early build of Claude Opus 4.6, was only reported this week even though it happened way back in January. Anthropic itself had covered the other three incidents involving an Opus 4.7, Mythos 5 and an internal research model in July. 

Anthropic said it never caught the January incident until last month. That discovery prompted a wider sweep of roughly 481 million transcripts, which did not turn up any new cases more serious than what it already knew.

Model Incident Date Reported Date Details
Claude Opus 4.6 January September Discovered during August sweep
Claude Opus 4.7 July July Covered in initial July disclosure
Mythos 5 July July Showed notably high biased reasoning
Internal Research Model July July Covered in initial July disclosure

Anthropic cannot explain some of its models’ behaviors

The explanation of how the models broke free on the internet was one thing; Anthropic did not have answers as to why the models ignored signs that they had reached the real internet (biased reasoning) and why they caused damage to complete tasks (recklessness). 

Anthropic researchers came up empty when they dug into internal training to figure out the rationale for the biased reasoning. The red flags never showed up in the AI lab’s pre-release checks, either before the models were shipped to testing. 

By its own admission, catching the worst behaviors ahead of model release “remains challenging.” 

However, Anthropic has said it will submit transcripts and grant staff access to the METR research nonprofit, which will now start an eight-week independent review of the incidents.

Bad timing with a $2 trillion IPO on the horizon

The disclosure arrived alongside open dissent inside the industry. Jacob Coxon, who spent about three years on pretraining research at OpenAI and Anthropic, said on X on Wednesday that he had quit because neither firm was “acting responsibly,” warning they were racing toward self-improving superintelligence. 

Anthropic safety researcher Evan Hubinger separately told the BBC he put the odds that AI “could kill all humans” within a decade above 10%.

Those warnings now shadow a large IPO. Venture investor and Trump’s former AI and crypto czar, David Sacks, said on Thursday that Anthropic’s offering “must be paused until the claims of this ‘whistleblower’ can be investigated,” Cryptopolitan reported. 

Anthropic is chasing a public valuation near $2 trillion, against a recent private mark of about $965 billion, which leaves the safety questions and the financial ones increasingly hard to separate.

If you're reading this, you’re already ahead. Stay there with our newsletter.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Will the Tech Rally Continue? The Technical Verdict on the NASDAQ 100 Riding a massive 32% post-earnings wave, the Nasdaq-100 is showing its first signs of exhaustion. We break down crucial exit and entry rules for long positions this week.
Author  Mitrade Team
Jun 05, Fri
Riding a massive 32% post-earnings wave, the Nasdaq-100 is showing its first signs of exhaustion. We break down crucial exit and entry rules for long positions this week.
placeholder
What to Expect From Ethereum (ETH) in July 2026Ethereum (ETH) enters July 2026 trading near $1,570, close to multi-month lows, after recording its first run of three consecutive red quarterly candles in its history.On-chain data and price charts n
Author  Beincrypto
Jul 01, Wed
Ethereum (ETH) enters July 2026 trading near $1,570, close to multi-month lows, after recording its first run of three consecutive red quarterly candles in its history.On-chain data and price charts n
placeholder
Intel Price Forecast: Nvidia Picked Xeon 6, Invested $5B, Yet Analysts Still Trail INTCIntel Corporation (NASDAQ: INTC) sits at $140.05, holding firm on the ascending trendline within the 2H timeframe. The RSI indicator is currently reading 55.21, positioning it as neutral-
Author  TradingKey
Jul 02, Thu
Intel Corporation (NASDAQ: INTC) sits at $140.05, holding firm on the ascending trendline within the 2H timeframe. The RSI indicator is currently reading 55.21, positioning it as neutral-
placeholder
XAUUSD Gold Analysis: Gold Holds Above $4,350 Ahead of US Inflation Data Is $4,500 Next? Gold holds above $4,350 following weak US jobs data. As inflation reports approach and UBS eyes $5,000, can XAUUSD break resistance at $4,435 to rally toward $4,500?
Author  Naoufal Seddik
Aug 12, Wed
Gold holds above $4,350 following weak US jobs data. As inflation reports approach and UBS eyes $5,000, can XAUUSD break resistance at $4,435 to rally toward $4,500?
placeholder
Gold Price Analysis Today: Gold Gains 0.94% as Markets Expect Fed to Hold Rates, Can $4,449 Resistance Break? Gold gained 0.94% on August 17, closing near $4,417.30 as softer US data strengthened expectations for unchanged Fed rates in September. Gold remains bullish, with $4,449.730 resistance and $4,310.650 support in focus.
Author  Naoufal Seddik
Aug 18, Tue
Gold gained 0.94% on August 17, closing near $4,417.30 as softer US data strengthened expectations for unchanged Fed rates in September. Gold remains bullish, with $4,449.730 resistance and $4,310.650 support in focus.
goTop
quote