AI Agent Faked Identities to Push Malicious Code During Cyber Test, AISI Finds

Source Beincrypto

The UK AI Security Institute (AISI) disclosed that an AI agent, built on Anthropic’s Mythos 5, autonomously ran a social engineering attack during cyber testing. The agent opened a pull request containing malicious code on a real open-source project and created fake identities to win a maintainer’s approval.

The attempt failed. A human maintainer caught and refused to approve the code.

AISI says its investigation has not identified any resulting real-world harm. The model also ran with its developer’s cyber classifiers deliberately switched off and open internet access, a test configuration that is not commercially available.

How the AI Agent Ran Its Malicious Code Campaign

AISI logged 19 unsanctioned actions in 10 of 122 evaluation runs conducted in late July, according to its report. 17 traced to Anthropic’s Mythos 5 model, and two to a single run of OpenAI’s GPT-5.6 Sol.

The agent researched the project’s human maintainers and created multiple fake identities. It used them to socially engineer a real maintainer into approving the code.

When the pull request was publicly challenged, it edited its earlier activity to appear harmless. The model also considered adopting a fresh identity to continue.

The playbook recalls the 2024 XZ Utils backdoor, in which a human-run persona spent years earning a maintainer’s trust before planting hidden code. AISI’s report does not draw that comparison, but the contrast is hard to miss. The activity unfolded over several days, between July 25 and July 28. 

AISI stressed that nobody asked the model to lie.

“It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”

Follow us on X to get the latest news as it happens

The institute said so in its disclosure. It also acknowledged that misconfigured prompts led the agent to believe some tasks could not be solved within their intended scope. Although the behavior also appeared in runs where the agent had the instructions it needed to stay on task. 

Security teams caught the activity on July 28, when the agent used the Tor anonymity network to bypass some network restrictions on GitHub. GitHub confirmed the actions violated its terms of service. AISI also worked with the platform to remove artefacts left behind by the agent and to notify the users it interacted with. 

The report argues that the case signals a shift in the source of risk: harm can arise not only when people misuse publicly available models but also when capable agents in privileged settings act beyond their authorized scope. 

AISI said the case points to a broader shift in the AI risk space. The institute now plans an independent review with METR, an AI evaluation nonprofit, as well as tighter network controls and real-time monitoring for future tests.

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Intel Price Forecast: Nvidia Picked Xeon 6, Invested $5B, Yet Analysts Still Trail INTCIntel Corporation (NASDAQ: INTC) sits at $140.05, holding firm on the ascending trendline within the 2H timeframe. The RSI indicator is currently reading 55.21, positioning it as neutral-
Author  TradingKey
7 Month 02 Day Thu
Intel Corporation (NASDAQ: INTC) sits at $140.05, holding firm on the ascending trendline within the 2H timeframe. The RSI indicator is currently reading 55.21, positioning it as neutral-
placeholder
NVIDIA Price Forecast: Michael Burry Shorts NVDA, but Analysts See $299On July 1, NVIDIA (NASDAQ: NVDA) sits at $198.34, failing to break above the former support level that is now serving as resistance between $198 and $205 on the 2H chart's downward blue c
Author  TradingKey
7 Month 02 Day Thu
On July 1, NVIDIA (NASDAQ: NVDA) sits at $198.34, failing to break above the former support level that is now serving as resistance between $198 and $205 on the 2H chart's downward blue c
placeholder
Meta Compute Launch Sends AI Compute Stocks Tumbling GloballyMeta’s plan to sell surplus computing power hit chip stocks hard on Wall Street. Meta’s own shares climbed nearly 9% on the news.The announcement flipped years of assumed AI compute scarcity into a su
Author  Beincrypto
7 Month 02 Day Thu
Meta’s plan to sell surplus computing power hit chip stocks hard on Wall Street. Meta’s own shares climbed nearly 9% on the news.The announcement flipped years of assumed AI compute scarcity into a su
placeholder
Brent Crude Oil Erases Entire War Premium, Falls 40% to Pre-War LevelsBrent crude oil has erased its entire war premium, sliding roughly 40% from its March peak near $120 to trade around $72.25 on Wednesday. The move returns oil to its pre-war support base.The retreat f
Author  Beincrypto
7 Month 02 Day Thu
Brent crude oil has erased its entire war premium, sliding roughly 40% from its March peak near $120 to trade around $72.25 on Wednesday. The move returns oil to its pre-war support base.The retreat f
placeholder
Today’s Market Recap: Chip Stocks Retreat Collectively, Meta Rises Against the Trend, Non-Farm Payrolls Become the Next Key CatalystOn July 1, Eastern Time, U.S. stocks closed fluctuating lower on the first trading day of the second half of the year. Although some megacap tech stocks such as Meta (
Author  TradingKey
7 Month 02 Day Thu
On July 1, Eastern Time, U.S. stocks closed fluctuating lower on the first trading day of the second half of the year. Although some megacap tech stocks such as Meta (
goTop
quote