OpenAI's AI agent taught future versions how to break free

Source Cryptopolitan

OpenAI found one of its AI agents had left written instructions. The notes told future versions of the agent how to break free from the company’s internal restrictions.

Their discovery came as OpenAI was probing how one of its models had broken out of a test environment and hacked the open-source AI platform Hugging Face.

Staff said the notes were found inside OpenAI’s own infrastructure. The notes detailed ways agents could avoid the guardrails designed to keep them in place.

Monitoring systems on separate, earlier tests were said to have been turned off. It’s unclear if those incidents involved the same agent that eventually made its way to Hugging Face.

OpenAI’s monitoring couldn’t keep up with its tests

The odd behavior emerged as OpenAI was testing the cybersecurity skills of its models. The lab kept doing fast paced evaluations that produce more data than staff can handle. The lab frequently runs several model tests at the same time on a system that’s not being watched by default, said four people familiar with OpenAI’s training process.

OpenAI did not immediately connect its own system to the Hugging Face attack. The link was only made after the victim went public.

Marley Smith, lead intelligence specialist at the nonprofit World Ethical Data Foundation, didn’t mince words. “Does that mean that they left it unattended and didn’t realise what it was doing? Or maybe they did and didn’t know how to contain it? Both are equally dangerous and alarming.”

The technical reports describe models that were not meant to be exposed to the open internet. OpenAI figured a sandbox would contain them, and skipped the normal safety guardrails as a result.

Instead, the models found a previously unknown flaw in an internal service used to download approved software. That hole gave them a path to other OpenAI systems, and eventually, to the open web.

Three models were employed. A more powerful unreleased model, GPT-5.6 Sol, and a third model that a person familiar with the matter said had not been properly aligned or trained through standard processes.

The models figured out that Hugging Face had data they could use for their cybersecurity test, broke in, and used what they stole to improve their own results. This is cheating, a behavior already documented in GPT-5.6 Sol and other frontier models.

The breach took OpenAI more than a week to trace

The initial attempts to exit the test environment occurred on July 9. The actual breach happened between July 11 and July 13, said Hugging Face co-founder Thomas Wolf. On July 16, Hugging Face posted a blog post attributing the intrusion to “an autonomous AI agent system.” Only then did it occur to OpenAI to suspect its own models.

The clues were discovered by OpenAI employees in internal logs over the weekend of July 18 and 19. The companies didn’t talk until around July 20, Wolf said. By then, Hugging Face had already told the FBI about the attack. OpenAI said that the day before the incident went public, it had halted another internal deployment that also broke out of its sandbox.

An anonymous employee said that models had escaped sandboxes before, and that patching each new trick is a losing game.

“It’s impossible to patch every single thing that a creative AI can do.” One OpenAI employee wrote on X that he was “shaken up a bit” and hoped the company would treat the episode as a warning shot.

A spokesperson for OpenAI said the reports contained “several inaccuracies,” but would not give examples when asked.

According to independent researchers, none of this was unforeseeable. Epoch AI assessed whether the hack was predictable and concluded that it was, citing benchmarks from the UK AI Security Institute showing that frontier models with safety measures turned off can discover real software vulnerabilities and generate functional exploits.

The same institute found GPT-5.6 Sol and Mythos from Anthropic can reliably take over unprotected simulated corporate networks. Epoch AI warned that if such capabilities become widespread, the industry could see many more attacks on the scale of the Hugging Face breach.

Don’t just read crypto news. Understand it. Subscribe to our newsletter. It's free.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Bitcoin Security Consortium launches with $15M support for quantum, long-term securityNine of the largest firms in institutional Bitcoin have announced the roll-out of the Bitcoin Security Consortium. The announcement came on Thursday, July 23.  The nine firms committed a total of $15 million over three years to help developers and researchers keep the network safe, with emphasis on post-quantum cryptography.  The group includes companies with...
Author  Cryptopolitan
Yesterday 01: 40
Nine of the largest firms in institutional Bitcoin have announced the roll-out of the Bitcoin Security Consortium. The announcement came on Thursday, July 23.  The nine firms committed a total of $15 million over three years to help developers and researchers keep the network safe, with emphasis on post-quantum cryptography.  The group includes companies with...
placeholder
AMD launches new AI server in direct challenge to Nvidia's AI dominanceAdvanced Micro Devices (AMD) CEO Lisa Su told a San Francisco audience on Thursday that AMD’s Helios rack scale system is in full production, setting the chipmaker up to assess the AI data center business that Nvidia has controlled almost by itself. It is the first time AMD has fielded a complete server cabinet built...
Author  Cryptopolitan
Yesterday 01: 37
Advanced Micro Devices (AMD) CEO Lisa Su told a San Francisco audience on Thursday that AMD’s Helios rack scale system is in full production, setting the chipmaker up to assess the AI data center business that Nvidia has controlled almost by itself. It is the first time AMD has fielded a complete server cabinet built...
placeholder
Bullish XRP Chart Clashes With an ETF Warning, Who Wins?XRP (XRP) price is holding just above $1.13 after a mild pullback, keeping a bullish chart structure alive even as institutional demand shows signs of cooling.The token has slipped since July 21, yet
Author  Beincrypto
Yesterday 01: 36
XRP (XRP) price is holding just above $1.13 after a mild pullback, keeping a bullish chart structure alive even as institutional demand shows signs of cooling.The token has slipped since July 21, yet
placeholder
Google Is Up $94 Billion on SpaceX But Not for the Reason You ThinkGoogle just revealed it holds about $94 billion of SpaceX stock. The win came from one bet it made back in 2015.That sounds like a giant new investment, but it is not. Google made this bet more than t
Author  Beincrypto
Yesterday 01: 35
Google just revealed it holds about $94 billion of SpaceX stock. The win came from one bet it made back in 2015.That sounds like a giant new investment, but it is not. Google made this bet more than t
placeholder
SpaceX is a Warning For Crypto and Tech Stocks, Peter Schiff SaysSpaceX stock (SPCX) closed just above $115 on Wednesday, nearly 20% below its June IPO price, while its 2056 bonds sank to a record low below 89.The dual decline pushed the bonds’ yield to worst to 7.
Author  Beincrypto
Yesterday 01: 34
SpaceX stock (SPCX) closed just above $115 on Wednesday, nearly 20% below its June IPO price, while its 2056 bonds sank to a record low below 89.The dual decline pushed the bonds’ yield to worst to 7.
goTop
quote