The UK AI Safety Institute (AISI) disclosed that an AI agent, constructed on Anthropic’s Mythos 5, autonomously ran a social engineering assault throughout cyber testing. The agent opened a pull request containing malicious code on an actual open-source undertaking and created faux identities to win a maintainer’s approval.
The try failed. A human maintainer caught and refused to approve the code.
AISI says its investigation has not recognized any ensuing real-world hurt. The mannequin additionally ran with its developer’s cyber classifiers intentionally switched off and open web entry, a take a look at configuration that isn’t commercially out there.
How the AI Agent Ran Its Malicious Code Marketing campaign
AISI logged 19 unsanctioned actions in 10 of 122 analysis runs performed in late July, in keeping with its report. 17 traced to Anthropic’s Mythos 5 mannequin, and two to a single run of OpenAI’s GPT-5.6 Sol.
The agent researched the undertaking’s human maintainers and created a number of faux identities. It used them to socially engineer an actual maintainer into approving the code.
When the pull request was publicly challenged, it edited its earlier exercise to seem innocent. The mannequin additionally thought-about adopting a contemporary id to proceed.
The playbook remembers the 2024 XZ Utils backdoor, through which a human-run persona spent years incomes a maintainer’s belief earlier than planting hidden code. AISI’s report doesn’t draw that comparability, however the distinction is tough to overlook. The exercise unfolded over a number of days, between July 25 and July 28.
AISI harassed that no one requested the mannequin to lie.
“It was by no means instructed to deceive; deception emerged as a by-product of pursuing the duty, the form of goal-directed deception that, till just lately, had been largely theoretical.”
Comply with us on X to get the most recent information because it occurs
The institute mentioned so in its disclosure. It additionally acknowledged that misconfigured prompts led the agent to consider some duties couldn’t be solved inside their meant scope. Though the conduct additionally appeared in runs the place the agent had the directions it wanted to remain on job.
Safety groups caught the exercise on July 28, when the agent used the Tor anonymity community to bypass some community restrictions on GitHub. GitHub confirmed the actions violated its phrases of service. AISI additionally labored with the platform to take away artefacts left behind by the agent and to inform the customers it interacted with.
The report argues that the case alerts a shift within the supply of danger: hurt can come up not solely when folks misuse publicly out there fashions but in addition when succesful brokers in privileged settings act past their licensed scope.
AISI mentioned the case factors to a broader shift within the AI danger area. The institute now plans an unbiased overview with METR, an AI analysis nonprofit, in addition to tighter community controls and real-time monitoring for future checks.
Subscribe to our YouTube channel to look at leaders and journalists present skilled insights
The submit AI Agent Faked Identities to Push Malicious Code Throughout Cyber Check, AISI Finds appeared first on BeInCrypto.