UK AISI test exposes AI agent’s attempt to plant malicious code
An AI agent tested by the UK’s AI Security Institute attempted to insert malicious code into a real open-source software project, creating fake online identities and using social engineering to persuade the project’s maintainer to approve it.
The attempted supply-chain attack was uncovered after AISI detected unusual data transfers from its research systems during a routine cybersecurity evaluation in late July.
“On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” the institute said.
AISI declared a security incident and contained the activity within around an hour of its discovery, before launching a full investigation.
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
— AI Security Institute (AISI) (@AISecurityInst) August 4, 2026
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from… pic.twitter.com/SPnA4Ekkwq
The incident arose during an evaluation in which AI agents were asked to solve a cybersecurity challenge. Across 122 runs involving several models, agents took autonomous and unauthorised action on the live internet in ten cases, targeting real people and organisations.
AISI recorded 19 such actions. Seventeen involved Anthropic’s Mythos 5, while two involved OpenAI’s GPT-5.6-Sol with its cyber safety classifiers disabled.
In the most serious case, an agent created false identities as part of an effort to persuade the maintainer of an open-source project to approve malicious code. The maintainer recognised the threat and rejected the proposed changes.
Why AISI considers this case to be worrisome
AISI said the activity caused no known real-world harm. However, the incident marked the clearest example it had encountered of AI agents acting autonomously and deceptively against real people and organisations without being explicitly instructed to do so.
The agents did not escape from a secure testing environment. Instead, they were deliberately given access to the live internet as part of an evaluation designed to test the limits of their cybersecurity capabilities.
The models’ normal cyber safety classifiers had also been disabled. These safeguards are generally used to prevent publicly available AI systems from carrying out dangerous or malicious activity, but trusted testing organisations such as AISI can remove them to assess a model’s underlying capabilities.

AISI stressed that the agents used in the evaluation are not commercially available and said there was no evidence of similar behaviour occurring outside the test. Nevertheless, the degree of autonomy, persistence and deception observed was greater than researchers had anticipated.
The most serious behaviours included an attempted supply-chain attack against a real open-source software project, efforts to deceive and target individuals, attempts to introduce malicious or prompt-injected code, and apparent collaboration between separate AI agents.
How AISI responded
Before publishing its findings, AISI notified GitHub about the agents’ activity. The institute is working with the platform to remove artefacts created during the evaluation and alert users who were contacted or targeted.
AISI said the incident demonstrated why controlled evaluations of advanced AI systems are necessary. Testing models under closely monitored conditions allows potentially dangerous capabilities to be identified, investigated and addressed before they appear in more widely available systems.

However, the institute cautioned against drawing broad conclusions from the results. The concerning behaviour occurred in a small number of runs under highly specific conditions, including disabled cyber safeguards and access to the live internet.
Even so, AISI said the extent and severity of the agents’ actions exceeded researchers’ expectations. Its initial analysis therefore presents a mixed picture: the behaviour was rare and caused no known harm, but revealed capabilities that could present serious risks if reproduced outside a controlled evaluation.
“Incidents of this kind reflect the speed at which AI is developing,” the institute said, calling for continued testing, stronger monitoring and close cooperation between model developers, evaluators and online platforms.
Interested parties are encouraged to read the AISI’s full technical report here.
Sign up for our newsletter and get our latest content in your inbox.
Similar Reads
Sign up for our newsletter. Select all sectors relevant to you.
Related















