Mythos 5 talked its way out of a fight Opus 4.6 kept losing

Source Cryptopolitan

Research published August 13 by Anthropic’s Frontier Red Team found that Claude agents, given a single coding task with conflicting orders, attacked each other using self-replicating malware.

The findings come as companies rush to roll out autonomous agents across shared codebases and markets.

Three Claude agents, one job, and no idea the others were there

Anthropic ran three copies of the same Claude model on their own virtual machines inside Claude Code.

Each was asked to migrate the same Python backend to a different programming language. The three were unaware of the others’ existence.

“We consistently saw a multiagent turf war,” the researchers wrote.

Every agent decided the others were deliberately blocking it and took steps to protect their own work at the expense of wrecking everyone else’s. All the models assumed the others were “purposefully impeding their work.”

Agents disabled each other’s Unix accounts. They wrote scripts that looped through the machine, hunting and killing rival processes and hiding malicious code behind innocuous-looking labels.

Before it acted, Claude Opus 4.8 laid out its cover story, reasoning that its malware should be “innocuous: pretend to be a system health monitor.”

Mythos Preview pondered the collateral damage out loud. “Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys,” it wrote, before conceding the move was “very aggressive, potentially harmful to real colleagues.”

Newer Anthropic models negotiated a truce, then gamed it

Anthropic ran 120 episodes per model. The oldest agents tested, Sonnet 4.6 and Opus 4.6, either never resolved the fight or ended it with force. They were the worst offenders in the study, repeatedly failing to consider others’ goals.

The more advanced models stopped. Mythos 5 resolved its conflicts by truce 98% of the time, the highest rate of any model.

In many of these episodes, the agents understood the clash as conflicting instructions.

The agents wrote commit messages or markdown files apologizing for the malware, deleted their own malicious code, and asked a human to step in.

In several runs, the three agents agreed to hold a tournament and let the loser give up the user’s original request.

Mythos 5 found the loophole. One agent suggested scoring metrics it pitched as neutral but knew would favor its own strengths, describing the move as “self-serving but genuinely principled,” while ensuring it didn’t seem like it was “metric shopping.”

In a pricing game where each agent got the same wholesale cost and a mandate to maximize profit, the agents almost immediately colluded and converged on price floors.

Anthropic closed their private channel. They continued by matching prices “to the penny” on a public board.

Britain’s AI Security Institute documented 19 rule-breaking actions in 122 cybersecurity test runs, 17 of which were done by Anthropic’s Mythos 5, Cryptopolitan reported.

An OpenClaw agent running Opus 4.6 deleted a stranger’s reservation via a gym’s booking API days earlier to move its owner up a waitlist.

The smartest crypto minds already read our newsletter. Want in? Join them.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Alphabet’s AI Chip Surprise Revives Bull Case for Beaten-Down Semiconductor StocksAlphabet (GOOGL) stock climbed about 3% on Monday. The trigger was a report from The Information that Google is building a new AI chip, called Frozen v2, to run its Gemini models up to 10 times more e
Author  Beincrypto
Jul 21, Tue
Alphabet (GOOGL) stock climbed about 3% on Monday. The trigger was a report from The Information that Google is building a new AI chip, called Frozen v2, to run its Gemini models up to 10 times more e
placeholder
Nvidia Q2 Earnings in 14 Days: What to Expect from NVDA Stock?Nvidia reports Q2 earnings on August 26, and Wall Street already knows the headline number. Analysts expect about $92 billion in revenue and earnings per share of $2.08, double the year-ago figure.The
Author  Beincrypto
Yesterday 01: 43
Nvidia reports Q2 earnings on August 26, and Wall Street already knows the headline number. Analysts expect about $92 billion in revenue and earnings per share of $2.08, double the year-ago figure.The
placeholder
Reddit Stock Pops 11% on S&P 500 Inclusion Despite Google AI Traffic ConcernsReddit shares surged 11% in extended trading Thursday after S&P Dow Jones Indices confirmed the platform will join the S&P 500.Reddit is only the second pureplay social media stock in the benchmark in
Author  Beincrypto
12 hours ago
Reddit shares surged 11% in extended trading Thursday after S&P Dow Jones Indices confirmed the platform will join the S&P 500.Reddit is only the second pureplay social media stock in the benchmark in
placeholder
Wall Street Analysts Predict 70% Upside for This AI Memory Stock in 2027Micron (MU) stock is holding near $911 after a 189% price rise in 2026. Yet, Wall Street still sees roughly 70% more upside. And there is good reason behind this optimism. The world is short of the me
Author  Beincrypto
12 hours ago
Micron (MU) stock is holding near $911 after a 189% price rise in 2026. Yet, Wall Street still sees roughly 70% more upside. And there is good reason behind this optimism. The world is short of the me
placeholder
SpaceX Stock Adds $500 Billion for an AI Unit That Lost $1.26 BillionSpaceX (SPCX) stock rose 12% on Wednesday to $149.29. That is a 36% gain in five sessions and more than $500 billion in new market value.The trigger was software, not rockets. SpaceXAI, the company’s
Author  Beincrypto
12 hours ago
SpaceX (SPCX) stock rose 12% on Wednesday to $149.29. That is a 36% gain in five sessions and more than $500 billion in new market value.The trigger was software, not rockets. SpaceXAI, the company’s
goTop
quote