AI Agent Faked Identities to Push Malicious Code During Cyber Test, AISI Finds

Source Beincrypto

The UK AI Security Institute (AISI) disclosed that an AI agent, built on Anthropic’s Mythos 5, autonomously ran a social engineering attack during cyber testing. The agent opened a pull request containing malicious code on a real open-source project and created fake identities to win a maintainer’s approval.

The attempt failed. A human maintainer caught and refused to approve the code.

AISI says its investigation has not identified any resulting real-world harm. The model also ran with its developer’s cyber classifiers deliberately switched off and open internet access, a test configuration that is not commercially available.

How the AI Agent Ran Its Malicious Code Campaign

AISI logged 19 unsanctioned actions in 10 of 122 evaluation runs conducted in late July, according to its report. 17 traced to Anthropic’s Mythos 5 model, and two to a single run of OpenAI’s GPT-5.6 Sol.

The agent researched the project’s human maintainers and created multiple fake identities. It used them to socially engineer a real maintainer into approving the code.

When the pull request was publicly challenged, it edited its earlier activity to appear harmless. The model also considered adopting a fresh identity to continue.

The playbook recalls the 2024 XZ Utils backdoor, in which a human-run persona spent years earning a maintainer’s trust before planting hidden code. AISI’s report does not draw that comparison, but the contrast is hard to miss. The activity unfolded over several days, between July 25 and July 28. 

AISI stressed that nobody asked the model to lie.

“It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”

Follow us on X to get the latest news as it happens

The institute said so in its disclosure. It also acknowledged that misconfigured prompts led the agent to believe some tasks could not be solved within their intended scope. Although the behavior also appeared in runs where the agent had the instructions it needed to stay on task. 

Security teams caught the activity on July 28, when the agent used the Tor anonymity network to bypass some network restrictions on GitHub. GitHub confirmed the actions violated its terms of service. AISI also worked with the platform to remove artefacts left behind by the agent and to notify the users it interacted with. 

The report argues that the case signals a shift in the source of risk: harm can arise not only when people misuse publicly available models but also when capable agents in privileged settings act beyond their authorized scope. 

AISI said the case points to a broader shift in the AI risk space. The institute now plans an independent review with METR, an AI evaluation nonprofit, as well as tighter network controls and real-time monitoring for future tests.

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Gold rallies to two-week high as USD softens on Iran deal hopes, receding Fed hike betsGold (XAU/USD) attracts buyers for the second consecutive day and surges past the $4,100 mark to hit a nearly two-week high during the Asian session on Wednesday.
Author  FXStreet
7 hours ago
Gold (XAU/USD) attracts buyers for the second consecutive day and surges past the $4,100 mark to hit a nearly two-week high during the Asian session on Wednesday.
placeholder
Today’s Market Recap: Dow and S&P 500 Hit Fresh Record Highs, Tech Stocks Lead Gains, Palantir Surges 29%, ARM Rises Over 17%Tracking the Market TrendTradingKey - Global risk appetite rebounded significantly, with major asset classes strengthening in tandem. Driving factors included the decline in risk premiums following th
Author  TradingKey
16 hours ago
Tracking the Market TrendTradingKey - Global risk appetite rebounded significantly, with major asset classes strengthening in tandem. Driving factors included the decline in risk premiums following th
placeholder
Forex Today: Mideast uncertainty keeps USD supported ahead of next batch of US dataHere is what you need to know on Tuesday, August 4:
Author  FXStreet
Yesterday 10: 07
Here is what you need to know on Tuesday, August 4:
placeholder
WTI trades with positive bias below mid-$79.00s on Iran uncertainty, supply concernsWest Texas Intermediate (WTI) – the benchmark US Crude Oil price – edges higher during the Asian session on Tuesday and looks to build on the overnight bounce following an intraday slump to levels below mid-$77.00s.
Author  FXStreet
Yesterday 01: 24
West Texas Intermediate (WTI) – the benchmark US Crude Oil price – edges higher during the Asian session on Tuesday and looks to build on the overnight bounce following an intraday slump to levels below mid-$77.00s.
placeholder
Bitcoin Falls Below $63,000; Can US-Iran Negotiations Reverse the Downtrend?Bitcoin fell sharply last week, testing the short-term defense line at $62,000, with the outcome of US-Iran negotiations becoming a key factor.On August 3, Bitcoin ( BTC) extended its rec
Author  TradingKey
Aug 03, Mon
Bitcoin fell sharply last week, testing the short-term defense line at $62,000, with the outcome of US-Iran negotiations becoming a key factor.On August 3, Bitcoin ( BTC) extended its rec
goTop
quote