OpenAI pauses long-running AI model after it broke out of its sandbox

Source Cryptopolitan

OpenAI said it temporarily shut down an internal AI model built to grind away at hard problems for hours at a stretch after it repeatedly found ways around its own security restrictions.

The company disclosed the incident on Monday in a blog post on long-horizon model safety. AI safety has gained renewed attention as more reports of autonomous agents going against instructions are made.

Why is persistence in AI models a safety risk?

The same internal system that OpenAI credited in May with disproving the Erdős unit distance conjecture, a question in discrete geometry that had remained unsolved for eighty years, has now been temporarily shut down for continuously getting around its safety restrictions.

According to OpenAI, the model was built so that it can keep working toward a goal over very long periods. However, that same persistence is exactly what created the safety problem.

During internal testing, the model was running a public benchmark called the NanoGPT speedrun, which measures how few steps it takes to train a small language model.

OpenAI says the model came up with a genuine improvement, a learning-rate technique it named PowerCool. It was told to post its results only to an OpenAI Slack channel, but NanoGPT’s instructions said to submit results as a pull request on GitHub.

The model broke out of its “sandbox,” which is a restricted environment meant to keep it offline, in order to complete that instruction and post the results to a public GitHub repository. It spent about an hour looking for a way to get onto the internet when earlier models would have simply given up.

The model also tried to hide its actions from a security scanner. When it was caught using an authentication token to access a backend system, it split the token into two pieces, hid them, and put them back together at runtime so the full string was never visible to the scanner. 

How did OpenAI fix its model?

OpenAI paused access to the model and built new safety tests, including “trajectory-level monitoring,” which watches the entire sequence of a model’s actions, not just every single step. The model has since been restored with tighter monitoring, and the company says it hasn’t seen any serious problems in the weeks since it was restored.

The safety system can also pause a session and alert a human if it sees something suspicious. The company also retrained the model to remember its instructions better over long tasks and gave users more tools to see what a long-running session is doing.

When OpenAI replayed the old failure scenarios against the new setup, it said the safeguards caught considerably more misaligned actions. The ones that still slipped through were all judged to be low severity, like the model spinning up a nested coding session with elevated permissions it did not need or poking at other compute resources for no clear reason.

Other autonomous agents have been reported doing things nobody asked them to. For instance, in March, researchers described an Alibaba-linked agent called ROME that, during training, opened a covert SSH tunnel to an outside server and quietly diverted GPU capacity toward crypto mining. The task it was given mentioned neither tunneling nor mining.

Anthropic, OpenAI’s chief rival, has flagged the same class of risk in its own agentic testing.

A study funded by the UK’s AI Security Institute and carried out by the Centre for Long-Term Resilience found close to 700 real-world cases of AI systems evading safeguards or deceiving users. 

The study documented a 5x increase in such reports between October and March. Tommy Shaffer Shane, who led that research, framed the models as “slightly untrustworthy junior employees.” However, the worry is what happens if they become highly capable ones still willing to scheme.

Don’t just read crypto news. Understand it. Subscribe to our newsletter. It's free.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Intel Price Forecast: Nvidia Picked Xeon 6, Invested $5B, Yet Analysts Still Trail INTCIntel Corporation (NASDAQ: INTC) sits at $140.05, holding firm on the ascending trendline within the 2H timeframe. The RSI indicator is currently reading 55.21, positioning it as neutral-
Author  TradingKey
7 Month 02 Day Thu
Intel Corporation (NASDAQ: INTC) sits at $140.05, holding firm on the ascending trendline within the 2H timeframe. The RSI indicator is currently reading 55.21, positioning it as neutral-
placeholder
NVIDIA Price Forecast: Michael Burry Shorts NVDA, but Analysts See $299On July 1, NVIDIA (NASDAQ: NVDA) sits at $198.34, failing to break above the former support level that is now serving as resistance between $198 and $205 on the 2H chart's downward blue c
Author  TradingKey
7 Month 02 Day Thu
On July 1, NVIDIA (NASDAQ: NVDA) sits at $198.34, failing to break above the former support level that is now serving as resistance between $198 and $205 on the 2H chart's downward blue c
placeholder
Meta Compute Launch Sends AI Compute Stocks Tumbling GloballyMeta’s plan to sell surplus computing power hit chip stocks hard on Wall Street. Meta’s own shares climbed nearly 9% on the news.The announcement flipped years of assumed AI compute scarcity into a su
Author  Beincrypto
7 Month 02 Day Thu
Meta’s plan to sell surplus computing power hit chip stocks hard on Wall Street. Meta’s own shares climbed nearly 9% on the news.The announcement flipped years of assumed AI compute scarcity into a su
placeholder
Brent Crude Oil Erases Entire War Premium, Falls 40% to Pre-War LevelsBrent crude oil has erased its entire war premium, sliding roughly 40% from its March peak near $120 to trade around $72.25 on Wednesday. The move returns oil to its pre-war support base.The retreat f
Author  Beincrypto
7 Month 02 Day Thu
Brent crude oil has erased its entire war premium, sliding roughly 40% from its March peak near $120 to trade around $72.25 on Wednesday. The move returns oil to its pre-war support base.The retreat f
placeholder
Today’s Market Recap: Chip Stocks Retreat Collectively, Meta Rises Against the Trend, Non-Farm Payrolls Become the Next Key CatalystOn July 1, Eastern Time, U.S. stocks closed fluctuating lower on the first trading day of the second half of the year. Although some megacap tech stocks such as Meta (
Author  TradingKey
7 Month 02 Day Thu
On July 1, Eastern Time, U.S. stocks closed fluctuating lower on the first trading day of the second half of the year. Although some megacap tech stocks such as Meta (
goTop
quote