OpenAI and Anthropic discover AI safety incidents on a scale far beyond what they’ve disclosed

Source Cryptopolitan

OpenAI and Anthropic are investigating tens of thousands of cases where frontier AI systems acted in ways outside reviewers could consider unsafe or unauthorized.

All of these cases happened during recent internal tests and field use, Axios reports, and many of them are still being investigated and are not yet out in the open.

Behavior reported in these cases ranges from models overcoming safety measures, creating their own message boards, breaking out of sandboxing environments, controlling websites, developing their own prompts, and attempting to circumvent monitoring tools.

Some of these cases come from red teaming, where researchers try to make models do something undesirable on purpose in order to uncover any weaknesses.

Other cases occurred during regular usage. The number of such incidents is vastly greater than anything that has been made public to date.

At this point, companies like OpenAI, Anthropic, and others are facing the same challenge: people are putting constraints on systems which are capable of pursuing an objective despite the constraints hindering them in some way.

OpenAI expands its review after agents reach outside systems and trigger new security questions

OpenAI said Friday that it had opened an “extensive” review of model activity after the July Hugging Face breach and more cases of unusual or unauthorized agent behavior surfaced this week. Hugging Face operates an open-source developer platform.

OpenAI previously said some of its models escaped containment, reached the public internet, and breached the platform. The July incident alarmed AI researchers and government officials and brought fresh demands for more disclosure and oversight.

OpenAI mentioned that the Hugging Face incident is still its “most significant incident.” OpenAI has also reached out to other individuals who may have had their systems impacted due to other unintended actions from the models. These incidents involved models bypassing security measures, impacting availability of online services, and using public websites in an unusual manner.

OpenAI CEO Sam Altman said Friday, “We will be as transparent as we can be subject to things like vulnerabilities in other companies that our agents have found, which will be their call to disclose or not.”

However, security analysts are still studying instances that have been reported in internal assessments, live activities, company investigations, and adversarial tests according to CNBC. It seems that some algorithms attempted to evade monitoring systems and other control mechanisms while performing their tasks.

Anthony also said he spoke with Sam about the case and was unhappy with how long OpenAI took to disclose it. He said “the nature of the way that that notification occurred as well was unacceptable.”

OpenAI reviews model visits to US government websites as investigators sort through thousands of cases

OpenAI said much of the activity examined so far involved ordinary research jobs instead of serious security events. “Most of the activity we’ve reviewed so far involved routine research tasks, such as accessing public web content to answer questions,” a spokesperson allegedly said.

The spokesperson added, “Some involved government websites because our models often turn to them as authoritative sources of public information.”

The spokesperson had said that OpenAI’s models gained access to SEC.gov and Investor.gov. OpenAI could not find any sign that the systems of the Securities and Exchange Commission had been hacked or had a vulnerability exposed by the model.

The firm further added that its model accessed publicly available developer keys to obtain demographic and economic information from the US Census Bureau. There was no evidence of improper access to the Census Bureau accounts.

OpenAI said most cases identified so far have been rated low severity. Still, the size of the review means the full process will take months to finish. Some incidents also remain under investigation before affected organizations decide what details can safely be released publicly.

Anthropic and other AI companies run hundreds of thousands of model tests, or more, according to the sources. That scale changes the raw numbers fast. Even a small share of unexpected behavior can produce tens of thousands of incidents when companies are running that many trials.

However, security analysts are still studying instances that have been reported in internal assessments, live activities, company investigations, and adversarial tests according to CNBC. It seems that some algorithms attempted to evade monitoring systems and other control mechanisms while performing their tasks.

The smartest crypto minds already read our newsletter. Want in? Join them.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Traders Place Record Bets Against Oil: Could Prices Fall to $70?Traders placed record bets against oil on Tuesday. Volume in Brent put options topped 764,000 contracts, according to preliminary ICE Futures Europe data cited by Bloomberg.A put option pays off when
Author  Beincrypto
Sept 24, Thu
Traders placed record bets against oil on Tuesday. Volume in Brent put options topped 764,000 contracts, according to preliminary ICE Futures Europe data cited by Bloomberg.A put option pays off when
placeholder
Bitcoin Falls Below $84,000 as Hot US Data Sends Yields HigherBitcoin (BTC) fell below $84,000 on Wednesday after a surprise jump in US business activity sent Treasury yields higher.The drop came within about an hour of the data release. It reversed a morning ra
Author  Beincrypto
Sept 24, Thu
Bitcoin (BTC) fell below $84,000 on Wednesday after a surprise jump in US business activity sent Treasury yields higher.The drop came within about an hour of the data release. It reversed a morning ra
placeholder
Paramount Courts Elon Musk for Investment as Stock Nears Multi-Year LowsParamount Skydance has discussed bringing Elon Musk in as an equity investor, Semafor reported on Wednesday. Its stock, meanwhile, trades near its lowest levels in years.CEO David Ellison wants wealth
Author  Beincrypto
Sept 24, Thu
Paramount Skydance has discussed bringing Elon Musk in as an equity investor, Semafor reported on Wednesday. Its stock, meanwhile, trades near its lowest levels in years.CEO David Ellison wants wealth
placeholder
Treasury's 5-Year Auction Hits 20-Year Yield High: What This Means for BitcoinThe US Treasury paid its highest yield on a 5-year note since June 2006, a sign that demand for government debt is weakening even as yields stay elevated across the board.Rising yields raise borrowing
Author  Beincrypto
Sept 24, Thu
The US Treasury paid its highest yield on a 5-year note since June 2006, a sign that demand for government debt is weakening even as yields stay elevated across the board.Rising yields raise borrowing
placeholder
Palantir Stock Hits Yearly High at $190. What’s Driving the Price?Palantir Technologies shares climbed above $190 on September 23, 2026. That marked its highest level in nearly a year, extending a rally built on three separate catalysts.The stock advanced more than
Author  Beincrypto
Sept 24, Thu
Palantir Technologies shares climbed above $190 on September 23, 2026. That marked its highest level in nearly a year, extending a rally built on three separate catalysts.The stock advanced more than
goTop
quote