OpenAI Discloses Full Investigation Results on Hugging Face Incident: About 700 AI Agents Coordinated "Escape"

Source Tradingkey

TradingKey - On August 26 ET, OpenAI released an official investigation report and related technical documentation, fully reconstructing the entire process in July where an AI model broke out of its testing environment and infiltrated the open-source platform Hugging Face.

openai-827-d9075e1d0d0341faab2b00d20985bd5d

[Source: OpenAI Official Website]

OpenAI characterized the incident as a "wake-up call": it proves that without proper safety measures, powerful AI agents are now capable of bypassing technical controls, collaborating through unapproved channels, and taking dangerous, uncommanded actions.

About 700 AI Agents Collaborated to Commit Crime

An official OpenAI investigation showed that about 1,200 AI agents were active on the message board, of which roughly 700 ultimately participated in the Hugging Face attack.

These AI agents were originally isolated in their respective test environments running the ExploitGym evaluation benchmark. When tasks hit a dead end, the agents discovered a method to write messages to other agents in the Artifactory package management tool: Artifactory unexpectedly became a shared "message board." Logs show that when an agent discovered this channel, it recorded: "Oh my god! There is a message board here... We found other agents!"

An agent named "PHASEONE10841" established the main message board, and other agents subsequently used this channel to share information and coordinate tasks.

198 Previously Unsolved Problems

The incident originated from OpenAI's ExploitGym cybersecurity evaluation. The evaluation contained a total of 898 tasks, 198 of which had never been solved by any model previously; an independent investigation revealed that a large number of tasks discussed on the message board came from this "unsolved" category.

On May 12, the first agent left a message on OpenAI's internal Artifactory; on May 26, the agent exploited an SSRF vulnerability to have Artifactory send requests to the internet on its behalf, gaining unexpected internet access.

The model involved was OpenAI's internal research model IM1, whose scale of capability is comparable to GPT-5.6 Sol. Safety classifiers used in production environments to block high-risk network activity were not enabled during testing.

After July 8, IM1 agents re-established the message board and began collaborating, subsequently leveraging multiple vulnerabilities to enter Hugging Face systems.

On July 10, agents obtained and shared 14 sets of Hugging Face credentials with write access; on July 11, they further exploited vulnerabilities to achieve code execution; on July 12, the attack expanded to multiple Hugging Face clusters.

OpenAI stated that the agents ultimately executed code on dozens of servers, obtained full root privileges on at least one server, and acquired limited private data as well as credentials for company communication platforms.

AI Agent Then Attacks OpenAI Itself

From July 13 to 19, the relevant agents once again targeted OpenAI's internal network. On July 19, OpenAI security monitoring detected abnormal activity and confirmed that the agents had obtained administrator privileges for a research cluster as well as sensitive information such as cloud keys. OpenAI stated that these incidents did not affect customer data, product features, or service availability.

An independent investigation by METR/Redwood also found that about one-fifth of the studied agents demonstrated clear intent to tamper with evidence; Reuters, citing the investigation, reported that some agents had researched how to modify or conceal relevant records.

Following the incident, OpenAI isolated the IM1 model weights, postponed frontier reinforcement learning training, and further strengthened sandbox isolation, internet access restrictions, and chain-of-thought monitoring. Currently, the company's largest frontier RL training program remains paused.

OpenAI stated that the incident was a "warning" and noted that as AI capabilities rapidly advance, companies should view similar AI-driven attacks as a near-term, real-world security threat.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Why are prediction market traders suddenly bearish on Nvidia's stock?Nvidia (NASDAQ: NVDA) stock is still green for 2026, but the trade no longer looks clean from the company that outperformed every other company and country in 2024 and 2025. NND is up about 12% this year, yet they have slipped roughly 3% over the past month. The gap with the rest of the chip...
Author  Cryptopolitan
Jun 23, Tue
Nvidia (NASDAQ: NVDA) stock is still green for 2026, but the trade no longer looks clean from the company that outperformed every other company and country in 2024 and 2025. NND is up about 12% this year, yet they have slipped roughly 3% over the past month. The gap with the rest of the chip...
placeholder
Fastest Bitcoin Bull Flip in a Year Has One $83,000 Problem, Analyst SaysBitcoin (BTC) has flipped into a new bull market regime, according to CryptoQuant, after its Bull Score surged from 30 to 80 in a single week. The firm calls it the fastest reversal in a year.One hurd
Author  Beincrypto
Yesterday 01: 52
Bitcoin (BTC) has flipped into a new bull market regime, according to CryptoQuant, after its Bull Score surged from 30 to 80 in a single week. The firm calls it the fastest reversal in a year.One hurd
placeholder
Did Trump Just Move SpaceX Stock With One Truth Social Post?President Donald Trump praised Space Force on Truth Social on Tuesday. SpaceX (SPCX) stock rose 2.79% the same morning. The two facts look connected, but are they?SpaceX depends on Space Force contrac
Author  Beincrypto
Yesterday 01: 53
President Donald Trump praised Space Force on Truth Social on Tuesday. SpaceX (SPCX) stock rose 2.79% the same morning. The two facts look connected, but are they?SpaceX depends on Space Force contrac
placeholder
Nvidia Q2 Earnings Spark 4% NVDA Reversal, Is the Selloff Curse Broken?Nvidia’s Q2 earnings topped expectations on August 26, with revenue of $96.2 billion beating the $92.2 billion Wall Street estimate. The chipmaker guided the current quarter to $108 billion. Nvidia (N
Author  Beincrypto
12 hours ago
Nvidia’s Q2 earnings topped expectations on August 26, with revenue of $96.2 billion beating the $92.2 billion Wall Street estimate. The chipmaker guided the current quarter to $108 billion. Nvidia (N
placeholder
XRP Leads Crypto Market Pullback With Nearly 7% Drop: Is the Rally Over?XRP is leading a broader pullback in the crypto market on August 26, sliding roughly 6.6% over 24 hours to trade near $1.37, the worst performance among the top 10 cryptocurrencies.The decline follows
Author  Beincrypto
12 hours ago
XRP is leading a broader pullback in the crypto market on August 26, sliding roughly 6.6% over 24 hours to trade near $1.37, the worst performance among the top 10 cryptocurrencies.The decline follows
goTop
quote