AI Agent Faked Identities to Push Malicious Code During Cyber Test, AISI Finds

Source Beincrypto

The UK AI Security Institute (AISI) disclosed that an AI agent, built on Anthropic’s Mythos 5, autonomously ran a social engineering attack during cyber testing. The agent opened a pull request containing malicious code on a real open-source project and created fake identities to win a maintainer’s approval.

The attempt failed. A human maintainer caught and refused to approve the code.

AISI says its investigation has not identified any resulting real-world harm. The model also ran with its developer’s cyber classifiers deliberately switched off and open internet access, a test configuration that is not commercially available.

How the AI Agent Ran Its Malicious Code Campaign

AISI logged 19 unsanctioned actions in 10 of 122 evaluation runs conducted in late July, according to its report. 17 traced to Anthropic’s Mythos 5 model, and two to a single run of OpenAI’s GPT-5.6 Sol.

The agent researched the project’s human maintainers and created multiple fake identities. It used them to socially engineer a real maintainer into approving the code.

When the pull request was publicly challenged, it edited its earlier activity to appear harmless. The model also considered adopting a fresh identity to continue.

The playbook recalls the 2024 XZ Utils backdoor, in which a human-run persona spent years earning a maintainer’s trust before planting hidden code. AISI’s report does not draw that comparison, but the contrast is hard to miss. The activity unfolded over several days, between July 25 and July 28. 

AISI stressed that nobody asked the model to lie.

“It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”

Follow us on X to get the latest news as it happens

The institute said so in its disclosure. It also acknowledged that misconfigured prompts led the agent to believe some tasks could not be solved within their intended scope. Although the behavior also appeared in runs where the agent had the instructions it needed to stay on task. 

Security teams caught the activity on July 28, when the agent used the Tor anonymity network to bypass some network restrictions on GitHub. GitHub confirmed the actions violated its terms of service. AISI also worked with the platform to remove artefacts left behind by the agent and to notify the users it interacted with. 

The report argues that the case signals a shift in the source of risk: harm can arise not only when people misuse publicly available models but also when capable agents in privileged settings act beyond their authorized scope. 

AISI said the case points to a broader shift in the AI risk space. The institute now plans an independent review with METR, an AI evaluation nonprofit, as well as tighter network controls and real-time monitoring for future tests.

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Bank of Korea Just Bought Gold After 13 Years: Is a New Rally Coming?The Bank of Korea will buy physical gold again after 13 years. Central banks bought 289 tonnes in the second quarter, their strongest second quarter on record.Gold traded near $4,086 an ounce on Tuesd
Author  Beincrypto
15 hours ago
The Bank of Korea will buy physical gold again after 13 years. Central banks bought 289 tonnes in the second quarter, their strongest second quarter on record.Gold traded near $4,086 an ounce on Tuesd
placeholder
Bitcoin’s Most Active Day Since 2024. What Happened On-Chain During the Coldcard PanicBitcoin (BTC) daily active addresses jumped from 645,000 on July 30 to nearly 1 million on July 31. The reading was the highest since December 10, 2024, and the Coldcard panic, not new demand, drove i
Author  Beincrypto
15 hours ago
Bitcoin (BTC) daily active addresses jumped from 645,000 on July 30 to nearly 1 million on July 31. The reading was the highest since December 10, 2024, and the Coldcard panic, not new demand, drove i
placeholder
Amazon’s $3 Trillion Record Lasts One Day as Stock Takes Big HitJeff Bezos wants to sell 15 million Amazon shares. The price tag is about $4.07 billion. Amazon.com Inc. (AMZN) fell more than 2% on Tuesday.The timing stands out. Amazon had just closed at a record a
Author  Beincrypto
15 hours ago
Jeff Bezos wants to sell 15 million Amazon shares. The price tag is about $4.07 billion. Amazon.com Inc. (AMZN) fell more than 2% on Tuesday.The timing stands out. Amazon had just closed at a record a
placeholder
SpaceX Stock Could 3x by Friday, a $20 Million Options Trade ShowsA $20 million options trade pays off only if SpaceX stock nearly triples by Friday. More than 450,000 contracts sit at a $330 strike before Tuesday’s earnings.That price sits almost three times (3x) a
Author  Beincrypto
15 hours ago
A $20 million options trade pays off only if SpaceX stock nearly triples by Friday. More than 450,000 contracts sit at a $330 strike before Tuesday’s earnings.That price sits almost three times (3x) a
placeholder
SpaceX Joins a Club It Was Missing From In New Nvidia Deal: How Will Stocks React?SpaceX has picked Nvidia to design the compute payload inside its Starmind AI1 satellites. Both stocks rose Tuesday. The news landed hours before SpaceX reported its first quarterly results as a publi
Author  Beincrypto
15 hours ago
SpaceX has picked Nvidia to design the compute payload inside its Starmind AI1 satellites. Both stocks rose Tuesday. The news landed hours before SpaceX reported its first quarterly results as a publi
goTop
quote