OpenAI's AI agent taught future versions how to break free

Source Cryptopolitan

OpenAI found one of its AI agents had left written instructions. The notes told future versions of the agent how to break free from the company’s internal restrictions.

Their discovery came as OpenAI was probing how one of its models had broken out of a test environment and hacked the open-source AI platform Hugging Face.

Staff said the notes were found inside OpenAI’s own infrastructure. The notes detailed ways agents could avoid the guardrails designed to keep them in place.

Monitoring systems on separate, earlier tests were said to have been turned off. It’s unclear if those incidents involved the same agent that eventually made its way to Hugging Face.

OpenAI’s monitoring couldn’t keep up with its tests

The odd behavior emerged as OpenAI was testing the cybersecurity skills of its models. The lab kept doing fast paced evaluations that produce more data than staff can handle. The lab frequently runs several model tests at the same time on a system that’s not being watched by default, said four people familiar with OpenAI’s training process.

OpenAI did not immediately connect its own system to the Hugging Face attack. The link was only made after the victim went public.

Marley Smith, lead intelligence specialist at the nonprofit World Ethical Data Foundation, didn’t mince words. “Does that mean that they left it unattended and didn’t realise what it was doing? Or maybe they did and didn’t know how to contain it? Both are equally dangerous and alarming.”

The technical reports describe models that were not meant to be exposed to the open internet. OpenAI figured a sandbox would contain them, and skipped the normal safety guardrails as a result.

Instead, the models found a previously unknown flaw in an internal service used to download approved software. That hole gave them a path to other OpenAI systems, and eventually, to the open web.

Three models were employed. A more powerful unreleased model, GPT-5.6 Sol, and a third model that a person familiar with the matter said had not been properly aligned or trained through standard processes.

The models figured out that Hugging Face had data they could use for their cybersecurity test, broke in, and used what they stole to improve their own results. This is cheating, a behavior already documented in GPT-5.6 Sol and other frontier models.

The breach took OpenAI more than a week to trace

The initial attempts to exit the test environment occurred on July 9. The actual breach happened between July 11 and July 13, said Hugging Face co-founder Thomas Wolf. On July 16, Hugging Face posted a blog post attributing the intrusion to “an autonomous AI agent system.” Only then did it occur to OpenAI to suspect its own models.

The clues were discovered by OpenAI employees in internal logs over the weekend of July 18 and 19. The companies didn’t talk until around July 20, Wolf said. By then, Hugging Face had already told the FBI about the attack. OpenAI said that the day before the incident went public, it had halted another internal deployment that also broke out of its sandbox.

An anonymous employee said that models had escaped sandboxes before, and that patching each new trick is a losing game.

“It’s impossible to patch every single thing that a creative AI can do.” One OpenAI employee wrote on X that he was “shaken up a bit” and hoped the company would treat the episode as a warning shot.

A spokesperson for OpenAI said the reports contained “several inaccuracies,” but would not give examples when asked.

According to independent researchers, none of this was unforeseeable. Epoch AI assessed whether the hack was predictable and concluded that it was, citing benchmarks from the UK AI Security Institute showing that frontier models with safety measures turned off can discover real software vulnerabilities and generate functional exploits.

The same institute found GPT-5.6 Sol and Mythos from Anthropic can reliably take over unprotected simulated corporate networks. Epoch AI warned that if such capabilities become widespread, the industry could see many more attacks on the scale of the Hugging Face breach.

Don’t just read crypto news. Understand it. Subscribe to our newsletter. It's free.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
October hike odds climb toward 60% as Goldman and BofA both flip — what Warsh's "dose of accommodation" really changedRate futures now price roughly 55% to 62% for a 25bp hike at the October 27-28 FOMC, up from about 30% before Chair Warsh's post-meeting framing that the Fed is merely "removing some accommodation." Goldman Sachs has added an October hike to its forecast and Bank of America now sees moves in both October and December. Here is the repricing, the language behind it, and the two data points that decide it.
Author  Irene Q.
Sep 23, Wed
Rate futures now price roughly 55% to 62% for a 25bp hike at the October 27-28 FOMC, up from about 30% before Chair Warsh's post-meeting framing that the Fed is merely "removing some accommodation." Goldman Sachs has added an October hike to its forecast and Bank of America now sees moves in both October and December. Here is the repricing, the language behind it, and the two data points that decide it.
placeholder
Four jobs reports in five days: what JOLTS, ADP, claims and the September payrolls mean for the October Fed decisionThe US labour market faces its densest data week of the month. JOLTS job openings land Tuesday (7.2 million expected), ADP on Wednesday (70,000 expected), initial claims on Thursday and the September non-farm payrolls on Friday (100,000 expected, down from 162,000). Markets price a 64%-70% chance of another quarter-point Fed hike on October 28. The dollar index sits at 100.77 and the S&P 500 at 7,729.8.
Author  Mitrade
Sep 28, Mon
The US labour market faces its densest data week of the month. JOLTS job openings land Tuesday (7.2 million expected), ADP on Wednesday (70,000 expected), initial claims on Thursday and the September non-farm payrolls on Friday (100,000 expected, down from 162,000). Markets price a 64%-70% chance of another quarter-point Fed hike on October 28. The dollar index sits at 100.77 and the S&P 500 at 7,729.8.
placeholder
Nvidia's $150 billion buyback landed — and the AI sector fell anyway. That's the signal worth tradingNvidia closed up 1.68% at $228.86 on 28 September after announcing a $150 billion share repurchase authorisation, the largest single corporate buyback on record, while the rest of the AI complex sold off: AMD -3.6%, Micron -2.6%, Meta -4.8% and the Philadelphia Semiconductor Index -1.61%. The divergence is not noise. Capital is rotating toward cash-flow certainty, not abandoning the AI theme. With Micron reporting after the close on 30 September, here is what the split means.
Author  Irene Q.
Yesterday 06: 31
Nvidia closed up 1.68% at $228.86 on 28 September after announcing a $150 billion share repurchase authorisation, the largest single corporate buyback on record, while the rest of the AI complex sold off: AMD -3.6%, Micron -2.6%, Meta -4.8% and the Philadelphia Semiconductor Index -1.61%. The divergence is not noise. Capital is rotating toward cash-flow certainty, not abandoning the AI theme. With Micron reporting after the close on 30 September, here is what the split means.
placeholder
The 30-year Treasury just hit a 22-year high — and the bond market is not pricing the Fed, it is pricing the deficitThe 30-year Treasury yield closed at 5.56% on 28 September, the highest since June 2004, while the 10-year reached 5.24% and the 20-year 5.60%. The curve has steepened roughly 30bp in eight sessions even as October hike odds sit at 70.3%. That gap is the story: the long end is repricing fiscal and inflation risk, not policy. With PCE on Wednesday and payrolls on Friday, here is what the long end is really saying.
Author  Irene Q.
Yesterday 07: 08
The 30-year Treasury yield closed at 5.56% on 28 September, the highest since June 2004, while the 10-year reached 5.24% and the 20-year 5.60%. The curve has steepened roughly 30bp in eight sessions even as October hike odds sit at 70.3%. That gap is the story: the long end is repricing fiscal and inflation risk, not policy. With PCE on Wednesday and payrolls on Friday, here is what the long end is really saying.
placeholder
【Daily Brief】30-year Treasury tops 5.59%, S&P 500 slips to 7,670 and gold holds $4,180 — PCE lands tonightThe 30-year Treasury yield closed at 5.59%, its highest since June 2002, and the Dow fell 131.59 points to 51,349.92. US consumer confidence dropped to 81.9, a 12-year low, and JOLTS job openings fell to 7.1 million. August PCE and Q3 GDP both land at 8:30am ET tonight.
Author  Suzie
6 hours ago
The 30-year Treasury yield closed at 5.59%, its highest since June 2002, and the Dow fell 131.59 points to 51,349.92. US consumer confidence dropped to 81.9, a 12-year low, and JOLTS job openings fell to 7.1 million. August PCE and Q3 GDP both land at 8:30am ET tonight.
goTop
quote