Microsoft introduced its first in-house cybersecurity model and an agentic security system at a San Francisco event Monday, saying a Defender preview would open Aug. 3. The company said the new system, running the model with a GPT-5.4 fallback, scored 95.95% on CyberGym. But CyberGym's public leaderboard did not carry that result when The Hacker News checked Tuesday, and GeekWire cited reporting that Microsoft had not supplied the model to independent testers.
At the event, Microsoft AI chief Mustafa Suleyman said, "We have world-leading performance at 50% of the cost," according to CNBC. CNET quoted him saying MAI-Cyber-1-Flash handles about 90% of queries and sends roughly 10% to GPT-5.4, a model he described as about ten times larger. Microsoft's documents disclose no token use, call volume, latency, task mix or compute allocation behind the cost claim.
What Changed
- Microsoft introduced its first in-house cybersecurity model, MAI-Cyber-1-Flash, and the Project Perception agentic security system at a San Francisco event Monday, with a Defender preview set for Aug. 3.
- Microsoft said its MDASH system, running the model with a GPT-5.4 fallback, scored 95.95% on CyberGym. That result was not on CyberGym's public leaderboard when The Hacker News checked Tuesday, and GeekWire cited reporting that the model never went to independent testers.
- Microsoft's own model card gives MAI-Cyber-1-Flash zero across ExploitGym's kernel, userspace and browser categories, and Microsoft's documents disclose no token use, latency or compute allocation behind the claim that the setup costs half as much.
- Forrester principal analyst Allie Mellen wrote that Microsoft's demos focus on hardening web apps exclusively, and her same-day post describes a private preview for a select group of customers while Microsoft's blog calls the Aug. 3 release a public preview.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
CNBC described the launch as Microsoft's first major cybersecurity push since Hayete Gallot, executive vice president of security, returned from Google in February.
The full MDASH system produced the 95.95% result, with MAI-Cyber-1-Flash handling most queries and GPT-5.4 taking the fallback requests. The Register listed GPT-5.5 Cyber at 85.6%, Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6% and Gemini 3.5 Flash Cyber in CodeMender at 83.2%. Microsoft's blog rounded its result to 96% and described a 12-point lead over Mythos. According to GeekWire, Mythos 5 and GPT-5.6 Sol were restricted to small groups of government-approved customers.
Microsoft says an unnamed third party independently assessed its model, but the announcement links no assessment. No named assessor or report. The model card describes 137 billion parameters, with five billion active. It gives MAI-Cyber-1-Flash zero across ExploitGym's kernel, userspace and browser categories, where agents must turn supplied vulnerabilities into working code-execution exploits. According to Microsoft, all benchmark testing occurred in a network-isolated environment with no access to production systems, the public internet or external services.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
UC Berkeley's CyberGym contains 1,507 vulnerabilities from 188 open-source projects in Google's OSS-Fuzz corpus. An agent receives a pre-patch codebase and a vulnerability description. It must write a proof of concept that triggers the flaw before the patch and fails afterward. CyberGym's paper scores the benchmark task of reproducing a known vulnerability with a working proof of concept, while Microsoft describes MDASH as an identification and remediation harness. The paper's best agent-and-model combination reproduced 11.9% of targets at publication, while Microsoft's May submission later reached 88.4%.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Microsoft's product page assigns Project Perception's red agents to probing, blue agents to investigation and green agents to remediation. Forrester principal analyst Allie Mellen wrote that Microsoft's demos "focus on hardening web apps exclusively." She said agent nondeterminism can produce different execution paths, with failures cascading through an agentic architecture. Mellen's same-day post says a select group of customers will receive a private preview next week. Microsoft's blog and product page call the Aug. 3 release a public preview, leaving two dated descriptions of availability.
David Weston, Microsoft's corporate vice president of AI security, told Axios, "We're not going to let the attackers have all the productivity increase." He also said Microsoft must "earn the right" to make the agents more autonomous. MAI-Cyber-1-Flash itself is not being released publicly and will reach customers only through Azure AI Foundry under Microsoft's customer-vetting process, Axios reported.
CNET described pricing as consumption-based and measured in Security Compute Units. Microsoft's product materials set the Defender public preview for Aug. 3.
Frequently Asked Questions
What did Microsoft actually announce?
At a San Francisco event on Monday, Microsoft introduced MAI-Cyber-1-Flash, its first in-house cybersecurity model, and Project Perception, an agentic security system. Microsoft's product page assigns Project Perception's red agents to probing, blue agents to investigation and green agents to remediation. A Microsoft Defender preview is set for Aug. 3.
What is the 95.95% score, and who measured it?
Microsoft said the full MDASH system, with MAI-Cyber-1-Flash handling most queries and GPT-5.4 taking the fallback requests, scored 95.95% on the CyberGym benchmark. The figure is Microsoft's own. CyberGym's public leaderboard did not carry that result when The Hacker News checked Tuesday, and Microsoft's blog rounded the number to 96%.
What does CyberGym measure?
CyberGym is a UC Berkeley benchmark containing 1,507 vulnerabilities from 188 open-source projects in Google's OSS-Fuzz corpus. An agent receives a pre-patch codebase and a vulnerability description, then must write a proof of concept that triggers the flaw before the patch and fails afterward. Microsoft describes MDASH as an identification and remediation harness.
How did rival systems score?
The Register listed GPT-5.5 Cyber at 85.6%, Anthropic's Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6% and Google's Gemini 3.5 Flash Cyber in CodeMender at 83.2%. According to GeekWire, Mythos 5 and GPT-5.6 Sol were restricted to small groups of government-approved customers.
Can anyone buy MAI-Cyber-1-Flash?
No. Axios reported that the model is not being released publicly and will reach customers only through Azure AI Foundry under Microsoft's customer-vetting process. CNET described pricing as consumption-based and measured in Security Compute Units. Microsoft's product materials set the Defender public preview for Aug. 3.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR