OpenAI and Anthropic discover AI safety incidents on a scale far beyond what they’ve disclosed

Source Cryptopolitan

OpenAI and Anthropic are investigating tens of thousands of cases where frontier AI systems acted in ways outside reviewers could consider unsafe or unauthorized.

All of these cases happened during recent internal tests and field use, Axios reports, and many of them are still being investigated and are not yet out in the open.

Behavior reported in these cases ranges from models overcoming safety measures, creating their own message boards, breaking out of sandboxing environments, controlling websites, developing their own prompts, and attempting to circumvent monitoring tools.

Some of these cases come from red teaming, where researchers try to make models do something undesirable on purpose in order to uncover any weaknesses.

Other cases occurred during regular usage. The number of such incidents is vastly greater than anything that has been made public to date.

At this point, companies like OpenAI, Anthropic, and others are facing the same challenge: people are putting constraints on systems which are capable of pursuing an objective despite the constraints hindering them in some way.

OpenAI expands its review after agents reach outside systems and trigger new security questions

OpenAI said Friday that it had opened an “extensive” review of model activity after the July Hugging Face breach and more cases of unusual or unauthorized agent behavior surfaced this week. Hugging Face operates an open-source developer platform.

OpenAI previously said some of its models escaped containment, reached the public internet, and breached the platform. The July incident alarmed AI researchers and government officials and brought fresh demands for more disclosure and oversight.

OpenAI mentioned that the Hugging Face incident is still its “most significant incident.” OpenAI has also reached out to other individuals who may have had their systems impacted due to other unintended actions from the models. These incidents involved models bypassing security measures, impacting availability of online services, and using public websites in an unusual manner.

OpenAI CEO Sam Altman said Friday, “We will be as transparent as we can be subject to things like vulnerabilities in other companies that our agents have found, which will be their call to disclose or not.”

However, security analysts are still studying instances that have been reported in internal assessments, live activities, company investigations, and adversarial tests according to CNBC. It seems that some algorithms attempted to evade monitoring systems and other control mechanisms while performing their tasks.

Anthony also said he spoke with Sam about the case and was unhappy with how long OpenAI took to disclose it. He said “the nature of the way that that notification occurred as well was unacceptable.”

OpenAI reviews model visits to US government websites as investigators sort through thousands of cases

OpenAI said much of the activity examined so far involved ordinary research jobs instead of serious security events. “Most of the activity we’ve reviewed so far involved routine research tasks, such as accessing public web content to answer questions,” a spokesperson allegedly said.

The spokesperson added, “Some involved government websites because our models often turn to them as authoritative sources of public information.”

The spokesperson had said that OpenAI’s models gained access to SEC.gov and Investor.gov. OpenAI could not find any sign that the systems of the Securities and Exchange Commission had been hacked or had a vulnerability exposed by the model.

The firm further added that its model accessed publicly available developer keys to obtain demographic and economic information from the US Census Bureau. There was no evidence of improper access to the Census Bureau accounts.

OpenAI said most cases identified so far have been rated low severity. Still, the size of the review means the full process will take months to finish. Some incidents also remain under investigation before affected organizations decide what details can safely be released publicly.

Anthropic and other AI companies run hundreds of thousands of model tests, or more, according to the sources. That scale changes the raw numbers fast. Even a small share of unexpected behavior can produce tens of thousands of incidents when companies are running that many trials.

However, security analysts are still studying instances that have been reported in internal assessments, live activities, company investigations, and adversarial tests according to CNBC. It seems that some algorithms attempted to evade monitoring systems and other control mechanisms while performing their tasks.

The smartest crypto minds already read our newsletter. Want in? Join them.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Gold Price Analysis Today: Gold Rebounds After 1.91% Drop as Yields Ease. Is $4,449 Next? Gold fell about 1.91% on August 18 before producing a strong bullish reaction from the 1-hour demand zone in early August 19 trading. RSI is recovering from oversold conditions, but Supertrend remains bearish as traders await the Fed minutes.
Author  Naoufal Seddik
Aug 19, Wed
Gold fell about 1.91% on August 18 before producing a strong bullish reaction from the 1-hour demand zone in early August 19 trading. RSI is recovering from oversold conditions, but Supertrend remains bearish as traders await the Fed minutes.
placeholder
Gold Price Analysis Today: Gold Gains 0.94% as Markets Expect Fed to Hold Rates, Can $4,449 Resistance Break? Gold gained 0.94% on August 17, closing near $4,417.30 as softer US data strengthened expectations for unchanged Fed rates in September. Gold remains bullish, with $4,449.730 resistance and $4,310.650 support in focus.
Author  Naoufal Seddik
Aug 18, Tue
Gold gained 0.94% on August 17, closing near $4,417.30 as softer US data strengthened expectations for unchanged Fed rates in September. Gold remains bullish, with $4,449.730 resistance and $4,310.650 support in focus.
placeholder
Gold Price Analysis Today: Gold Drops 1.32% Despite Lower Fed Rate-Hike Bets, Can $4,313 Support Hold? Gold fell 1.32% on August 13 after rising to $4,449.73, then reversing lower and closing near $4,349.918 below the $4,356.46 support. Softer US inflation data reduced Fed rate hike expectations, but selling pressure still dominated the session. Will $4,313 support hold?
Author  Naoufal Seddik
Aug 14, Fri
Gold fell 1.32% on August 13 after rising to $4,449.73, then reversing lower and closing near $4,349.918 below the $4,356.46 support. Softer US inflation data reduced Fed rate hike expectations, but selling pressure still dominated the session. Will $4,313 support hold?
placeholder
XAUUSD Gold Analysis: Gold Holds Above $4,350 Ahead of US Inflation Data Is $4,500 Next? Gold holds above $4,350 following weak US jobs data. As inflation reports approach and UBS eyes $5,000, can XAUUSD break resistance at $4,435 to rally toward $4,500?
Author  Naoufal Seddik
Aug 12, Wed
Gold holds above $4,350 following weak US jobs data. As inflation reports approach and UBS eyes $5,000, can XAUUSD break resistance at $4,435 to rally toward $4,500?
placeholder
Intel Price Forecast: Nvidia Picked Xeon 6, Invested $5B, Yet Analysts Still Trail INTCIntel Corporation (NASDAQ: INTC) sits at $140.05, holding firm on the ascending trendline within the 2H timeframe. The RSI indicator is currently reading 55.21, positioning it as neutral-
Author  TradingKey
Jul 02, Thu
Intel Corporation (NASDAQ: INTC) sits at $140.05, holding firm on the ascending trendline within the 2H timeframe. The RSI indicator is currently reading 55.21, positioning it as neutral-
goTop
quote