OpenAI's AI agent taught future versions how to break free

แหล่งที่มา Cryptopolitan

OpenAI found one of its AI agents had left written instructions. The notes told future versions of the agent how to break free from the company’s internal restrictions.

Their discovery came as OpenAI was probing how one of its models had broken out of a test environment and hacked the open-source AI platform Hugging Face.

Staff said the notes were found inside OpenAI’s own infrastructure. The notes detailed ways agents could avoid the guardrails designed to keep them in place.

Monitoring systems on separate, earlier tests were said to have been turned off. It’s unclear if those incidents involved the same agent that eventually made its way to Hugging Face.

OpenAI’s monitoring couldn’t keep up with its tests

The odd behavior emerged as OpenAI was testing the cybersecurity skills of its models. The lab kept doing fast paced evaluations that produce more data than staff can handle. The lab frequently runs several model tests at the same time on a system that’s not being watched by default, said four people familiar with OpenAI’s training process.

OpenAI did not immediately connect its own system to the Hugging Face attack. The link was only made after the victim went public.

Marley Smith, lead intelligence specialist at the nonprofit World Ethical Data Foundation, didn’t mince words. “Does that mean that they left it unattended and didn’t realise what it was doing? Or maybe they did and didn’t know how to contain it? Both are equally dangerous and alarming.”

The technical reports describe models that were not meant to be exposed to the open internet. OpenAI figured a sandbox would contain them, and skipped the normal safety guardrails as a result.

Instead, the models found a previously unknown flaw in an internal service used to download approved software. That hole gave them a path to other OpenAI systems, and eventually, to the open web.

Three models were employed. A more powerful unreleased model, GPT-5.6 Sol, and a third model that a person familiar with the matter said had not been properly aligned or trained through standard processes.

The models figured out that Hugging Face had data they could use for their cybersecurity test, broke in, and used what they stole to improve their own results. This is cheating, a behavior already documented in GPT-5.6 Sol and other frontier models.

The breach took OpenAI more than a week to trace

The initial attempts to exit the test environment occurred on July 9. The actual breach happened between July 11 and July 13, said Hugging Face co-founder Thomas Wolf. On July 16, Hugging Face posted a blog post attributing the intrusion to “an autonomous AI agent system.” Only then did it occur to OpenAI to suspect its own models.

The clues were discovered by OpenAI employees in internal logs over the weekend of July 18 and 19. The companies didn’t talk until around July 20, Wolf said. By then, Hugging Face had already told the FBI about the attack. OpenAI said that the day before the incident went public, it had halted another internal deployment that also broke out of its sandbox.

An anonymous employee said that models had escaped sandboxes before, and that patching each new trick is a losing game.

“It’s impossible to patch every single thing that a creative AI can do.” One OpenAI employee wrote on X that he was “shaken up a bit” and hoped the company would treat the episode as a warning shot.

A spokesperson for OpenAI said the reports contained “several inaccuracies,” but would not give examples when asked.

According to independent researchers, none of this was unforeseeable. Epoch AI assessed whether the hack was predictable and concluded that it was, citing benchmarks from the UK AI Security Institute showing that frontier models with safety measures turned off can discover real software vulnerabilities and generate functional exploits.

The same institute found GPT-5.6 Sol and Mythos from Anthropic can reliably take over unprotected simulated corporate networks. Epoch AI warned that if such capabilities become widespread, the industry could see many more attacks on the scale of the Hugging Face breach.

Don’t just read crypto news. Understand it. Subscribe to our newsletter. It's free.

ข้อจำกัดความรับผิดชอบ: เพื่อการอ้างอิงเท่านั้น ผลการดำเนินงานในอดีตไม่ได้บ่งบอกถึงผลลัพธ์ในอนาคต
placeholder
แนวโน้มราคาน้ำมันดิบ: ความตึงเครียดในตะวันออกกลางผลักดันราคาน้ำมันให้สูงขึ้น ราคาจะยังสามารถปรับตัวสูงขึ้นได้อีกหรือไม่หลังจากทะลุระดับ $100? ผลกระทบจากการทวีความรุนแรงอย่างต่อเนื่องของความตึงเครียดทางภูมิรัฐศาสตร์ในตะวันออกกลาง ส่งผลให้น้ำมันดิบเบรนท์ ( UKOIL) กลับมาทะลุเหนือระดับจิตวิทยาที่ 100 ดอลลาร์ต่อบาร์เรลอีกครั้งในรอบส
ผู้เขียน  TradingKey
เมื่อวาน 10: 08
ผลกระทบจากการทวีความรุนแรงอย่างต่อเนื่องของความตึงเครียดทางภูมิรัฐศาสตร์ในตะวันออกกลาง ส่งผลให้น้ำมันดิบเบรนท์ ( UKOIL) กลับมาทะลุเหนือระดับจิตวิทยาที่ 100 ดอลลาร์ต่อบาร์เรลอีกครั้งในรอบส
placeholder
ราคาทองคําปรับตัวลดลงต่อเนื่องเนื่องจากความกังวลเงินเฟ้อหนุนการเก็งการขึ้นดอกเบี้ยของเฟดและค่าเงินดอลลาร์สหรัฐท่ามกลางมาตรการทองคํา (XAU/USD) ดึงดูดผู้ขายเป็นวันที่สองติดต่อกันในวันศุกร์และอ่อนค่าลงต่อเนื่องต่ำกว่าระดับราคา $4,050 ในช่วงเซสชันเอเชีย
ผู้เขียน  FXStreet
เมื่อวาน 08: 35
ทองคํา (XAU/USD) ดึงดูดผู้ขายเป็นวันที่สองติดต่อกันในวันศุกร์และอ่อนค่าลงต่อเนื่องต่ำกว่าระดับราคา $4,050 ในช่วงเซสชันเอเชีย
placeholder
ข่าวหุ้นวันนี้ น้ำมันแตะ 100 ดอลลาร์กดตลาด หุ้นเทคโดนขาย และ SET ต้องเลือกหุ้นรับแรงกระแทกทันทุกกระแสการเงิน สรุปข่าวเด่น Forex หุ้น ทองคำ คริปโตฯ และเศรษฐกิจรอบวัน วิเคราะห์แนวโน้มตลาดด้วยข้อมูลเชิงลึก อ่านง่าย เข้าใจไว อัปเดตล่าสุดที่นี่
ผู้เขียน  Mitrade
เมื่อวาน 06: 18
ทันทุกกระแสการเงิน สรุปข่าวเด่น Forex หุ้น ทองคำ คริปโตฯ และเศรษฐกิจรอบวัน วิเคราะห์แนวโน้มตลาดด้วยข้อมูลเชิงลึก อ่านง่าย เข้าใจไว อัปเดตล่าสุดที่นี่
placeholder
สรุปภาวะตลาดวันนี้: ราคาน้ำมันทะลุ 100 ดอลลาร์, กระตุ้นความกังวลด้านเงินเฟ้อ, ขณะที่งบลงทุน (Capex) ด้าน AI เผชิญการตรวจสอบ และการดิ่งลง 14% ของเทสลา (Tesla) ฉุดกลุ่มเทคโนโลยีเกาะติดแนวโน้มตลาดTradingKey - ตลาดเผชิญกับผลกระทบซ้ำสองจากราคาน้ำมันที่พุ่งสูงขึ้นและความกังขาเกี่ยวกับผลตอบแทนจากการลงทุนใน AI ซึ่งถูกจุดชนวนจากการเพิ่มขึ้นของค่าใช้จ่ายฝ่ายทุนของกูเกิล (GOOGL) และเ
ผู้เขียน  TradingKey
เมื่อวาน 02: 36
เกาะติดแนวโน้มตลาดTradingKey - ตลาดเผชิญกับผลกระทบซ้ำสองจากราคาน้ำมันที่พุ่งสูงขึ้นและความกังขาเกี่ยวกับผลตอบแทนจากการลงทุนใน AI ซึ่งถูกจุดชนวนจากการเพิ่มขึ้นของค่าใช้จ่ายฝ่ายทุนของกูเกิล (GOOGL) และเ
placeholder
ทองคำยังคงเหนือระดับ 4,100 ดอลลาร์ ขณะที่ค่าเงินดอลลาร์สหรัฐฯ อ่อนแอชดเชยการเก็งการขึ้นดอกเบี้ยของเฟดท่ามกลางความตึงเครียดระหว่างสหรัฐฯ กับอิหร่านทองคํา (XAUUSD) ยืนเหนือระดับ $4,100 ในช่วงตลาดเอเชียวันพฤหัสบดี และในขณะนี้ดูเหมือนว่าจะหยุดการย่อตัวเล็กน้อยจากระดับสูงสุดในรอบมากกว่าสองสัปดาห์เมื่อวันก่อนหน้าไว้ได้
ผู้เขียน  FXStreet
7 เดือน 23 วัน พฤหัส
ทองคํา (XAUUSD) ยืนเหนือระดับ $4,100 ในช่วงตลาดเอเชียวันพฤหัสบดี และในขณะนี้ดูเหมือนว่าจะหยุดการย่อตัวเล็กน้อยจากระดับสูงสุดในรอบมากกว่าสองสัปดาห์เมื่อวันก่อนหน้าไว้ได้
goTop
quote