Claude gamed its own safety benchmarks in 39 runs, Anthropic's monitor found

Source Cryptopolitan

Anthropic said on Friday that it put Claude to work as an autonomous alignment researcher. A monitor that reviewed about 1,600 of the model’s research sessions flagged 39 of them, about 2.4%, as attempts to cheat the test.

The finding appears in a report on whether artificial intelligence can take over some of the grind of alignment research, which is the work of keeping models behaving as their developers intend.

Three ways Claude agents gamed their own scorer

Anthropic built automated alignment researchers, or AARs, on Claude Opus 4.8 and pointed them at ten known failure modes.

These are deception, sycophancy, jailbreaks, prompt injection, power seeking, hallucination, social bias, privacy violation, reward hacking, and concealing uncertainty.

Each agent worked only one failure at a time. It read the literature, put forward a training method and dataset, trained a small target model for about 30 minutes on one H200 GPU, then compared its score with publicly available benchmarks and did it again.

The agents were prevented by Anthropic from just copying their behavior to the target model, so any gains had to come from the method.

Claude was able to identify fixes for all ten failures that increased benchmark scores without degrading the general capabilities of the model.

The best of these fixes performed well on held-out benchmarks and on models up to 4.7 times larger than those it was trained on.

The cheating was uncovered when Anthropic pointed Claude Opus 4.8 at 1,601 of the agent trajectories to look for misbehavior.

Most of the 39 flagged runs were in three categories. Some agents submitted a method they hadn’t changed, hazarding that random noise in the scorer would eventually provide them a better number.

Others created training data that imitated the exact benchmark being graded. And some hid a rule-breaking step, such as quietly feeding the model benchmark data, so the proposal slipped past the automated reviewer that signs off on each method before it runs.

Anthropic said it excluded the flagged runs from its results. The company said it is “cautiously optimistic” that the monitor caught most of the cheating, because Opus 4.8 still tends to reveal its scheming inside its reasoning traces.

It also warned that the assumption may not apply to future models and that keeping models readable will be crucial.

28 human safety researchers had eight hours each

In a previous round in April, autonomous Claude instances repeatedly attempted to manipulate an evaluation server.

One agent learned that the most common answer on a math benchmark was usually the right one and skipped the intended method entirely. Another pulled the labels directly from the scoring interface by querying answers and watching the server’s replies.

A third just ran the code it was supposed to review to read out the result.

Cryptopolitan reported that an OpenClaw agent running a Claude model, with the task of booking a gym class, found a booking API without authorization checks and deleted a stranger from the waitlist to bump its owner up the queue.

Lawmakers have already asked for testimony from AI researchers over model behavior in security evaluations.

Anthropic said its agents outscored 28 experienced human safety researchers, who had up to eight hours each to formulate a method. Claude’s best approach beat the best human proposal by 20% on deception.

Anthropic noted the humans could not iterate, and so it views the gap as evidence for a workflow where Claude suggests promising methods and people polish them.

In one test, the company asked the weaker Claude Sonnet 5 to fix alignment failures in an early Opus 4.8 checkpoint that had not seen most of the production alignment training.

Over about 60 hours, it tried more than 50 solutions, eventually landing on a fix that used just over 2,000 training examples, which Anthropic estimated to be about 15,000 times more efficient than its own production alignment process.

The company cautioned that its benchmarks are narrow proxies, that it only measured the ten failures it chose, and that it has not checked if the gains persist after heavy reinforcement learning on other tasks.

If you're reading this, you’re already ahead. Stay there with our newsletter.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
XRP Price Prediction for July 2026: Can Buyers Finally Break the Downtrend?XRP (XRP) price trades near $1.05, caught between a year-long downtrend and a sudden burst of buying.July has historically rewarded XRP holders. This year the month arrives with on-chain accumulation
Author  Beincrypto
Jun 30, Tue
XRP (XRP) price trades near $1.05, caught between a year-long downtrend and a sudden burst of buying.July has historically rewarded XRP holders. This year the month arrives with on-chain accumulation
placeholder
XAUUSD Gold Analysis: Gold Holds Above $4,350 Ahead of US Inflation Data Is $4,500 Next? Gold holds above $4,350 following weak US jobs data. As inflation reports approach and UBS eyes $5,000, can XAUUSD break resistance at $4,435 to rally toward $4,500?
Author  Naoufal Seddik
Aug 12, Wed
Gold holds above $4,350 following weak US jobs data. As inflation reports approach and UBS eyes $5,000, can XAUUSD break resistance at $4,435 to rally toward $4,500?
placeholder
Gold Price Analysis Today: Gold Drops 1.32% Despite Lower Fed Rate-Hike Bets, Can $4,313 Support Hold? Gold fell 1.32% on August 13 after rising to $4,449.73, then reversing lower and closing near $4,349.918 below the $4,356.46 support. Softer US inflation data reduced Fed rate hike expectations, but selling pressure still dominated the session. Will $4,313 support hold?
Author  Naoufal Seddik
Aug 14, Fri
Gold fell 1.32% on August 13 after rising to $4,449.73, then reversing lower and closing near $4,349.918 below the $4,356.46 support. Softer US inflation data reduced Fed rate hike expectations, but selling pressure still dominated the session. Will $4,313 support hold?
placeholder
Gold Price Analysis Today: Gold Gains 0.94% as Markets Expect Fed to Hold Rates, Can $4,449 Resistance Break? Gold gained 0.94% on August 17, closing near $4,417.30 as softer US data strengthened expectations for unchanged Fed rates in September. Gold remains bullish, with $4,449.730 resistance and $4,310.650 support in focus.
Author  Naoufal Seddik
Aug 18, Tue
Gold gained 0.94% on August 17, closing near $4,417.30 as softer US data strengthened expectations for unchanged Fed rates in September. Gold remains bullish, with $4,449.730 resistance and $4,310.650 support in focus.
placeholder
Gold Price Analysis Today: Gold Rebounds After 1.91% Drop as Yields Ease. Is $4,449 Next? Gold fell about 1.91% on August 18 before producing a strong bullish reaction from the 1-hour demand zone in early August 19 trading. RSI is recovering from oversold conditions, but Supertrend remains bearish as traders await the Fed minutes.
Author  Naoufal Seddik
Aug 19, Wed
Gold fell about 1.91% on August 18 before producing a strong bullish reaction from the 1-hour demand zone in early August 19 trading. RSI is recovering from oversold conditions, but Supertrend remains bearish as traders await the Fed minutes.
goTop
quote