OpenAI pauses long-running AI model after it broke out of its sandbox

Source Cryptopolitan

OpenAI said it temporarily shut down an internal AI model built to grind away at hard problems for hours at a stretch after it repeatedly found ways around its own security restrictions.

The company disclosed the incident on Monday in a blog post on long-horizon model safety. AI safety has gained renewed attention as more reports of autonomous agents going against instructions are made.

Why is persistence in AI models a safety risk?

The same internal system that OpenAI credited in May with disproving the Erdős unit distance conjecture, a question in discrete geometry that had remained unsolved for eighty years, has now been temporarily shut down for continuously getting around its safety restrictions.

According to OpenAI, the model was built so that it can keep working toward a goal over very long periods. However, that same persistence is exactly what created the safety problem.

During internal testing, the model was running a public benchmark called the NanoGPT speedrun, which measures how few steps it takes to train a small language model.

OpenAI says the model came up with a genuine improvement, a learning-rate technique it named PowerCool. It was told to post its results only to an OpenAI Slack channel, but NanoGPT’s instructions said to submit results as a pull request on GitHub.

The model broke out of its “sandbox,” which is a restricted environment meant to keep it offline, in order to complete that instruction and post the results to a public GitHub repository. It spent about an hour looking for a way to get onto the internet when earlier models would have simply given up.

The model also tried to hide its actions from a security scanner. When it was caught using an authentication token to access a backend system, it split the token into two pieces, hid them, and put them back together at runtime so the full string was never visible to the scanner. 

How did OpenAI fix its model?

OpenAI paused access to the model and built new safety tests, including “trajectory-level monitoring,” which watches the entire sequence of a model’s actions, not just every single step. The model has since been restored with tighter monitoring, and the company says it hasn’t seen any serious problems in the weeks since it was restored.

The safety system can also pause a session and alert a human if it sees something suspicious. The company also retrained the model to remember its instructions better over long tasks and gave users more tools to see what a long-running session is doing.

When OpenAI replayed the old failure scenarios against the new setup, it said the safeguards caught considerably more misaligned actions. The ones that still slipped through were all judged to be low severity, like the model spinning up a nested coding session with elevated permissions it did not need or poking at other compute resources for no clear reason.

Other autonomous agents have been reported doing things nobody asked them to. For instance, in March, researchers described an Alibaba-linked agent called ROME that, during training, opened a covert SSH tunnel to an outside server and quietly diverted GPU capacity toward crypto mining. The task it was given mentioned neither tunneling nor mining.

Anthropic, OpenAI’s chief rival, has flagged the same class of risk in its own agentic testing.

A study funded by the UK’s AI Security Institute and carried out by the Centre for Long-Term Resilience found close to 700 real-world cases of AI systems evading safeguards or deceiving users. 

The study documented a 5x increase in such reports between October and March. Tommy Shaffer Shane, who led that research, framed the models as “slightly untrustworthy junior employees.” However, the worry is what happens if they become highly capable ones still willing to scheme.

Don’t just read crypto news. Understand it. Subscribe to our newsletter. It's free.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Alibaba jumps 5.4% after unveiling flagship AI modelAlibaba claims its latest AI model is the world’s strongest after Anthropic’s top system. This move comes as China now stands face to face with U.S AI rivals. The company unveiled a preview of Qwen3.8-Max, the flagship model in its Qwen lineup, through an official X post on Sunday. The preview is now accessible on...
Author  Cryptopolitan
15 hours ago
Alibaba claims its latest AI model is the world’s strongest after Anthropic’s top system. This move comes as China now stands face to face with U.S AI rivals. The company unveiled a preview of Qwen3.8-Max, the flagship model in its Qwen lineup, through an official X post on Sunday. The preview is now accessible on...
placeholder
Microsoft to deploy AMD's new Helios AI system in its data centers, lifting AMD's stock over 3%AMD’s stock increased by almost 5% on Monday following the chip company’s announcement that it and Microsoft were strengthening their business partnership, with Microsoft agreeing to integrate AMD’s new rack-scale AI technology into its data centers. Investors now have more motivation to support AMD as a legitimate rival since the move is perceived as a...
Author  Cryptopolitan
15 hours ago
AMD’s stock increased by almost 5% on Monday following the chip company’s announcement that it and Microsoft were strengthening their business partnership, with Microsoft agreeing to integrate AMD’s new rack-scale AI technology into its data centers. Investors now have more motivation to support AMD as a legitimate rival since the move is perceived as a...
placeholder
Apple Stock Price Prediction: Can July Earnings Push AAPL Past $5 Trillion?Apple stock (AAPL) is within roughly 4% of a $5 Trillion milestone after a rapid rally. The next earnings report will test whether fundamentals can support the move.Apple shares currently remain 9.8%
Author  Beincrypto
15 hours ago
Apple stock (AAPL) is within roughly 4% of a $5 Trillion milestone after a rapid rally. The next earnings report will test whether fundamentals can support the move.Apple shares currently remain 9.8%
placeholder
Alphabet’s AI Chip Surprise Revives Bull Case for Beaten-Down Semiconductor StocksAlphabet (GOOGL) stock climbed about 3% on Monday. The trigger was a report from The Information that Google is building a new AI chip, called Frozen v2, to run its Gemini models up to 10 times more e
Author  Beincrypto
15 hours ago
Alphabet (GOOGL) stock climbed about 3% on Monday. The trigger was a report from The Information that Google is building a new AI chip, called Frozen v2, to run its Gemini models up to 10 times more e
placeholder
Bitcoin Reclaims $65,000 as BTC ETF Inflows Return: Is the Worst Over?US spot Bitcoin (BTC) exchange-traded funds (ETFs) pulled in $75.7 million last week, their second winning week in a row. Bitcoin also reclaimed $65,000 on Monday as hopes grew that US-Iran talks may
Author  Beincrypto
15 hours ago
US spot Bitcoin (BTC) exchange-traded funds (ETFs) pulled in $75.7 million last week, their second winning week in a row. Bitcoin also reclaimed $65,000 on Monday as hopes grew that US-Iran talks may
goTop
quote