In short
- Microsoft says MDASH scored 95.95% on CyberGym, topping GPT-5.5 Cyber, Mythos 5, GPT-5.6 Sol, and Gemini 3.5 Flash Cyber.
- MAI-Cyber-1-Flash handles as much as 90% of the workload, whereas MDASH sends the toughest circumstances to GPT-5.4.
- The scanner is in personal preview by means of Microsoft Defender, the place groups can evaluation findings and generate proposed fixes.
Microsoft has launched its first devoted cybersecurity mannequin named MAI-Cyber-1-Flash and plugged it into MDASH, a vulnerability-hunting system that it says beats Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol whereas costing 50% lower than Microsoft’s present greatest MDASH configuration, in line with the corporate.
The mixed setup scored 95.95% on CyberGym, in line with Microsoft. CyberGym is a benchmark that asks AI brokers to breed 1,507 identified vulnerabilities throughout 188 open-source tasks, then scores them by the share efficiently reproduced in a managed surroundings.
That put MDASH forward of GPT-5.5 Cyber at 85.6%, Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6%, and Gemini 3.5 Flash Cyber at 83.2%. The result’s self-reported by Microsoft and had not appeared on CyberGym’s public leaderboard at publication time, although the benchmark makes use of a public check set and an outlined success metric.
MAI-Cyber-1-Flash doesn’t work alone. Microsoft says it handles as much as 90% of duties, whereas MDASH routes the toughest 10% to GPT-5.4. That issues as a result of tokens—the chunks of textual content an AI processes—value cash each time a mannequin reads code or produces a solution.

That is the primary time a Microsoft mannequin constructed for effectivity is able to beating a dense state-of-the-art mannequin constructed with basic capabilities in thoughts. “When mixed with MDASH, (MAI-Cyber-1-Flash) delivers world-class efficiency at 50 % of the price of main fashions,” Microsoft CEO Satya Nadella wrote.
As we speak, we’re asserting a collection of updates that give clients frontier-grade safety at half the associated fee.
MAI-Cyber-1-Flash is our first cybersecurity mannequin, constructed floor as much as discover probably the most difficult vulnerabilities in advanced code bases. When mixed with MDASH, it… pic.twitter.com/npcIihN1H7
— Satya Nadella (@satyanadella) July 27, 2026
A mannequin (on this case MAI-Cyber-1-Flash) is the AI that causes over the code. A harness (MDASH on this case) is the equipment round it: the brokers, instruments, checks, and workflow that determine the place to look, problem suspected findings, take away duplicates, and show {that a} bug may be triggered.
MDASH makes use of greater than 100 specialised brokers assigned to audit code, debate whether or not a discovering is real, and construct a proof of idea—a working demonstration that the flaw exists.
Microsoft stated occasional scans and delayed patches have gotten out of date as AI makes bug discovery cheaper. The corporate argues that a long time of safety knowledge give it a bonus, including, “Nobody can manufacture this historical past.”
Ever because the launch of Claude Mythos, cybersecurity consultants have been attempting to beat or match its capabilities. Researchers reproduced Mythos-style vulnerability looking with public fashions for below $30 per scan. Dawid Moczadło, one of many researchers concerned, stated “the moat is shifting from mannequin entry to validation.”
The scarce half is turning into the system that proves findings with out burying builders below false alarms.
Decrypt additionally reported that GPT-5.5 Cyber had just lately taken the general public CyberGym lead with an 85.6% rating, narrowly beating Mythos. Microsoft’s rating is about 10 factors above GPT-5.5 Cyber and seven.5 factors above MDASH’s earlier consequence, nevertheless it evaluates the complete system fairly than MAI-Cyber-1-Flash by itself.
Microsoft is placing MDASH into personal preview by means of Microsoft Safety Publicity Administration within the Defender portal. Prospects can scan Git repositories, see findings ranked from unlikely to confirmed, and use the Defender CLI to generate proposed code fixes for developer evaluation.
The preview presently limits repositories to roughly 256MB and permits one concurrent scan per tenant. Venture Notion is predicted to increase the identical multi-agent strategy past code scanning into broader risk monitoring and remediation workflows.
Every day Debrief Publication
Begin every single day with the highest information tales proper now, plus unique options, a podcast, movies and extra.
