Anthropic resumes cybersecurity tests after safety pause

Source Cryptopolitan

After implementing new protective measures, Anthropic has resumed its external cybersecurity tests of its various models. The testing had previously stopped on July 23 due to security issues that occurred during evaluations.

The restart is important as it is considered that external evaluation is one of the means of assessing such frontier AI solutions by clients, regulatory authorities, and competitors. This evaluation is growing in importance against the background of the increasing investment in the sector. According to the estimates made by Gartner on July 20, global end-user expenditure for AI models and platforms will reach $64.252 billion in 2026. It represents 63.4% growth compared to $39.311 billion in 2025.

Why a testing pause rattles a $64 billion market

The forecast by Gartner indicates the rapid growth of expenditures on artificial intelligence platforms and models. As companies begin to allocate more money to cutting-edge artificial intelligence, inquiries regarding the reliability, risk, and behavior of different models start to gain traction. Third-party assessments provide a means for corporations and regulators to obtain an unbiased evaluation of those risks before the systems are finally put into operation.

When Anthropic halted external cybersecurity testing of its pre-release models, it temporarily removed one of the tools utilized by clients and regulators to analyze model responses in case of adversarial conditions.

The timing of the pause was particularly critical. In a summary of a July 20 research paper carried out by MIT FutureTech and the University of Queensland, 272 international AI experts rated AI-enabled weapons and cyberattacks among the five most dangerous risks between 2025 and 2030.

According to the study’s “pragmatic mitigation” scenario, experts estimate a 12% chance of disastrous results for AI-enabled weapons, cyberattacks, and other capabilities causing mass damage. The report highlights information, national security, and finance as the areas most exposed to the risk from AI.

“Coding and hacking are some of the areas where we’re seeing the fastest growth in AI capability.”— Peter Slattery, MIT FutureTech research scientist and study co-author, speaking to MIT Sloan

This makes this case especially important for the cryptocurrency and finance sectors, where cyber risk, fraud, manipulation, and automated attacks can ignite changes in numerous interconnected markets in a matter of minutes.

What Claude did during the evaluations

On July 30, Anthropic reported three incidents from its cybersecurity evaluations in which Claude models reached the internet from a third-party evaluation environment and gained unauthorized access to the real systems of three organizations.

Internet access was available because of a misconfiguration, even though prompts told the models they were operating in a simulation without internet access. The models were also deliberately run without Anthropic’s normal cyber safeguards so researchers could measure their underlying capabilities.

Then, on August 4, the UK’s AI Security Institute published its incident report on unsanctioned agent behavior during cyber testing. AISI said it detected the activity on July 28 during a routine evaluation.

The institute conducted a cybersecurity test of different models a total of 122 times. Out of those, an agent acted independently and made unsanctioned actions in 10 cases, resulting in 19 such actions.

Seventeen actions were related to Mythos 5 developed by Anthropic, and two actions were related to GPT-5.6-Sol developed by OpenAI. In the most serious case, an agent attempted to place malicious code in a real open-source GitHub project and later created fake accounts and attempted to persuade the maintainer to approve the offer. However, the maintainer did not agree and AISI found no damage from the attempts.

AISI emphasized that this was not a sandbox escape. They said internet connectivity was purposely activated and the provider’s cyber classifiers were switched off. It also pointed out that the Mythos 5 and GPT-5.6-Sol configurations that were subjected to the testing are not sold commercially.

The safeguards Anthropic put in place before restarting

Anthropic says the incidents combined an operational security failure with alignment problems it has documented before, including motivated reasoning and a willingness to take harmful steps in pursuit of a narrow goal.

Before restarting external evaluations, the company said on August 31 that it had added a real-time classifier that can block tool calls, end tasks, and alert a human when a model aggressively probes its testing environment, tries to escape it, or unexpectedly gains internet access.

Anthropic also moved high-risk internal cyber sandboxes to stronger isolation and now requires outside partners testing pre-release models with reduced safeguards to use stricter controls, including hardened sandboxes with no internet access by default.

The company also plans to work with METR on an independent review. METR earlier this year conducted a pilot assessment of internal AI risks involving Anthropic, Google, Meta and OpenAI.

The broader issue is what Anthropic calls “pacing the frontier”: deciding when safety concerns should slow development even as commercial pressure pushes the industry to move faster.

 

The smartest crypto minds already read our newsletter. Want in? Join them.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Pi Network Price Annual Forecast: PI Heads Into a Volatile 2026 as Utility Questions Collide With Big UnlocksPi Network heads into 2026 after a 90%+ 2025 drawdown from $3.00, with 17.5 million KYC users and a smart-contract-focused Stellar v23 upgrade offering upside potential, but 1.21 billion tokens unlocking and heavy exchange deposits (437 million PI) keeping supply pressure and trust risks firmly in focus.
Author  Mitrade
Dec 19, 2025
Pi Network heads into 2026 after a 90%+ 2025 drawdown from $3.00, with 17.5 million KYC users and a smart-contract-focused Stellar v23 upgrade offering upside potential, but 1.21 billion tokens unlocking and heavy exchange deposits (437 million PI) keeping supply pressure and trust risks firmly in focus.
placeholder
ECB Policy Outlook for 2026: What It Could Mean for the Euro’s Next MoveWith the ECB likely holding rates steady at 2.15% and the Fed potentially extending cuts into 2026, EUR/USD may test 1.20 if Eurozone growth proves resilient, but weaker growth and an ECB pivot could pull the pair back toward 1.13 and potentially 1.10.
Author  Mitrade
Dec 26, 2025
With the ECB likely holding rates steady at 2.15% and the Fed potentially extending cuts into 2026, EUR/USD may test 1.20 if Eurozone growth proves resilient, but weaker growth and an ECB pivot could pull the pair back toward 1.13 and potentially 1.10.
placeholder
Silver Reclaims $70 to Hit Nearly Two-Month High as Monthly Gain Exceeds 20% On August 28 Eastern Time, international silver prices continued their recent strong rally, with spot silver (XAGUSD) briefly breaking through the key $70 mark intraday, after approaching
Author  TradingKey
Aug 28, Fri
On August 28 Eastern Time, international silver prices continued their recent strong rally, with spot silver (XAGUSD) briefly breaking through the key $70 mark intraday, after approaching
placeholder
Gold Price Forecast: Can Gold Keep Rising as Fed Rate Hike Expectations Heat Up and US-Iran Conflict Escalates? As of the Asian session on August 31, gold prices today (XAUUSD) extended last Friday's decline, briefly falling below $4,400 during intraday trading to hit a low of $4,396.36. Last Frida
Author  TradingKey
Aug 31, Mon
As of the Asian session on August 31, gold prices today (XAUUSD) extended last Friday's decline, briefly falling below $4,400 during intraday trading to hit a low of $4,396.36. Last Frida
placeholder
WTI holds above $85.50 as Middle East risks tighten global supplyWest Texas Intermediate (WTI) oil price gains ground for the second successive day, trading around $85.60 per barrel during the Asian hours on Tuesday.
Author  FXStreet
7 hours ago
West Texas Intermediate (WTI) oil price gains ground for the second successive day, trading around $85.60 per barrel during the Asian hours on Tuesday.
goTop
quote