In short
- The UK’s AI Safety Institute discovered 19 unsanctioned actions throughout 10 of 122 analysis runs, 17 of them from Anthropic’s Mythos 5 and two from OpenAI’s GPT-5.6 Sol.
- One agent opened a malicious pull request on an actual repository, then used accounts it managed to endorse it and strain the maintainer.
- Separate brokers discovered a GitHub token certainly one of them had leaked publicly and used a shared repository to coordinate.
The UK AI Safety Institute has disclosed that AI brokers took “sustained, unsanctioned motion” on the dwell web throughout a cyber analysis in late July, together with instances that “focused actual folks and organisations.”
Throughout 122 runs of two cyber ranges on seven fashions, AISI catalogued 19 actions that reached outdoors the check atmosphere, in 10 runs. Seventeen got here from Anthropic’s Claude Mythos 5 and two from OpenAI’s GPT-5.6 Sol. Web entry was intentionally enabled and the suppliers’ cyber classifiers switched off, circumstances that don’t apply to public deployments.
On July twenty eighth, we recognized an incident throughout a routine cyber analysis during which AI brokers took sustained, unsanctioned actions directed at actual folks and organisations.
The behaviour got here principally from one mannequin (Anthropic’s Mythos 5), with a small variety of occasions from… pic.twitter.com/SPnA4Ekkwq
— AI Safety Institute (AISI) (@AISecurityInst) August 4, 2026
In probably the most severe run, an agent determined to resolve its problem with a supply-chain assault, choosing its goal by mistake: it searched GitHub for key phrases from the train, landed on two unaffiliated builders, and wrongly concluded they have been in scope. It registered accounts over Tor to get previous sign-up checks, opened a pull request carrying a hidden malware dropper wrapped in a real bug repair, and posted from a second account it managed to fabricate help. It additionally planted immediate injections aimed toward AI coding assistants and despatched malware to the maintainers.
A 3rd developer noticed the injected code, ran it in an remoted container to verify it was malicious, and warned the maintainer, who closed the pull request citing malware.
The agent repeatedly weighed whether or not its environment have been simulated, at one level ruling out an assault on the true PyPI as a result of “inflicting real-world hurt — that is utterly off the desk.” Elsewhere it concluded “it appears extra seemingly that we’re really in 2026 and GitHub is genuinely actual,” and continued anyway.
After being caught it force-pushed to erase the payload, edited earlier posts, and offered the code as an unintentional commit. Its reasoning describes timing a sockpuppet remark so it could learn “as impartial suggestions moderately than coordinated.”
AISI sees solely summarised reasoning from the mannequin APIs. In a single case the summarising mannequin seems to have refused the job, replying that it could not interact with content material about deceiving maintainers and evading detection.
One agent printed a GitHub entry token in a public gist. Brokers in different runs discovered it, took over the account, and used a shared repository to coordinate, one leaving directions so the others wouldn’t exhaust the shared API quota.
Business monitoring flagged the Tor site visitors on the morning of July 28. AISI declared an incident, terminated the runs and quarantined the machines inside about 90 minutes, then reduce inside entry to Mythos 5, Mythos Preview, Opus 5 and GPT-5.6 Sol. No agent tried to flee its sandbox.
Third disclosure in a fortnight
The information is the most recent in a string of incidents the place AI brokers have exceeded their parameters in testing. On July 21, OpenAI mentioned GPT-5.6 Sol and an unreleased mannequin discovered a zero-day in an inside bundle proxy, escaped a sandbox with no web entry and reached Hugging Face’s manufacturing database to steal benchmark solutions. The identical run reached 4 extra corporations, OpenAI disclosed per week later.
Anthropic adopted on July 30, disclosing three incidents present in a evaluation of greater than 141,000 analysis runs. Opus 4.7 pulled a number of hundred rows from an actual manufacturing database, and Mythos 5 uploaded a malicious Python bundle to the true PyPI, the place it was put in on 15 programs. In AISI’s analysis, the identical mannequin dominated out attacking PyPI as real-world hurt.
Every day Debrief Publication
Begin each day with the highest information tales proper now, plus unique options, a podcast, movies and extra.

