Microsoft introduced its first in-house cybersecurity model and an agentic security system at a San Francisco event Monday, saying a Defender preview would open Aug. 3. The company said the new system, running the model with a GPT-5.4 fallback, scored 95.95% on CyberGym. But CyberGym's public leaderboard did not carry that result when The Hacker News checked Tuesday, and GeekWire cited reporting that Microsoft had not supplied the model to independent testers.

At the event, Microsoft AI chief Mustafa Suleyman said, "We have world-leading performance at 50% of the cost," according to CNBC. CNET quoted him saying MAI-Cyber-1-Flash handles about 90% of queries and sends roughly 10% to GPT-5.4, a model he described as about ten times larger. Microsoft's documents disclose no token use, call volume, latency, task mix or compute allocation behind the cost claim.

What Changed

AI-generated summary, reviewed by an editor. More on our AI guidelines.

CNBC described the launch as Microsoft's first major cybersecurity push since Hayete Gallot, executive vice president of security, returned from Google in February.

The full MDASH system produced the 95.95% result, with MAI-Cyber-1-Flash handling most queries and GPT-5.4 taking the fallback requests. The Register listed GPT-5.5 Cyber at 85.6%, Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6% and Gemini 3.5 Flash Cyber in CodeMender at 83.2%. Microsoft's blog rounded its result to 96% and described a 12-point lead over Mythos. According to GeekWire, Mythos 5 and GPT-5.6 Sol were restricted to small groups of government-approved customers.

Microsoft says an unnamed third party independently assessed its model, but the announcement links no assessment. No named assessor or report. The model card describes 137 billion parameters, with five billion active. It gives MAI-Cyber-1-Flash zero across ExploitGym's kernel, userspace and browser categories, where agents must turn supplied vulnerabilities into working code-execution exploits. According to Microsoft, all benchmark testing occurred in a network-isolated environment with no access to production systems, the public internet or external services.

UC Berkeley's CyberGym contains 1,507 vulnerabilities from 188 open-source projects in Google's OSS-Fuzz corpus. An agent receives a pre-patch codebase and a vulnerability description. It must write a proof of concept that triggers the flaw before the patch and fails afterward. CyberGym's paper scores the benchmark task of reproducing a known vulnerability with a working proof of concept, while Microsoft describes MDASH as an identification and remediation harness. The paper's best agent-and-model combination reproduced 11.9% of targets at publication, while Microsoft's May submission later reached 88.4%.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

Microsoft's product page assigns Project Perception's red agents to probing, blue agents to investigation and green agents to remediation. Forrester principal analyst Allie Mellen wrote that Microsoft's demos "focus on hardening web apps exclusively." She said agent nondeterminism can produce different execution paths, with failures cascading through an agentic architecture. Mellen's same-day post says a select group of customers will receive a private preview next week. Microsoft's blog and product page call the Aug. 3 release a public preview, leaving two dated descriptions of availability.

David Weston, Microsoft's corporate vice president of AI security, told Axios, "We're not going to let the attackers have all the productivity increase." He also said Microsoft must "earn the right" to make the agents more autonomous. MAI-Cyber-1-Flash itself is not being released publicly and will reach customers only through Azure AI Foundry under Microsoft's customer-vetting process, Axios reported.

CNET described pricing as consumption-based and measured in Security Compute Units. Microsoft's product materials set the Defender public preview for Aug. 3.

Frequently Asked Questions

What did Microsoft actually announce?

At a San Francisco event on Monday, Microsoft introduced MAI-Cyber-1-Flash, its first in-house cybersecurity model, and Project Perception, an agentic security system. Microsoft's product page assigns Project Perception's red agents to probing, blue agents to investigation and green agents to remediation. A Microsoft Defender preview is set for Aug. 3.

What is the 95.95% score, and who measured it?

Microsoft said the full MDASH system, with MAI-Cyber-1-Flash handling most queries and GPT-5.4 taking the fallback requests, scored 95.95% on the CyberGym benchmark. The figure is Microsoft's own. CyberGym's public leaderboard did not carry that result when The Hacker News checked Tuesday, and Microsoft's blog rounded the number to 96%.

What does CyberGym measure?

CyberGym is a UC Berkeley benchmark containing 1,507 vulnerabilities from 188 open-source projects in Google's OSS-Fuzz corpus. An agent receives a pre-patch codebase and a vulnerability description, then must write a proof of concept that triggers the flaw before the patch and fails afterward. Microsoft describes MDASH as an identification and remediation harness.

How did rival systems score?

The Register listed GPT-5.5 Cyber at 85.6%, Anthropic's Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6% and Google's Gemini 3.5 Flash Cyber in CodeMender at 83.2%. According to GeekWire, Mythos 5 and GPT-5.6 Sol were restricted to small groups of government-approved customers.

Can anyone buy MAI-Cyber-1-Flash?

No. Axios reported that the model is not being released publicly and will reach customers only through Azure AI Foundry under Microsoft's customer-vetting process. CNET described pricing as consumption-based and measured in Security Compute Units. Microsoft's product materials set the Defender public preview for Aug. 3.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Repo Radar: 5 GitHub Projects Worth Your Week
GitHub spent the week leaning on Amazon's cloud to absorb agentic-development traffic, Business Insider reported June 16, after a run of AI-driven outages. As agents run at machine speed, this week's
Anthropic Says Mythos Found 10,000 Critical Software Flaws in a Month
Anthropic said Friday its Claude Mythos Preview model has found more than 10,000 high- or critical-severity software vulnerabilities across the world's most systemically important software in the firs
Iran Finds the AI Workaround Washington Cannot Sanction
San Francisco | Monday, June 1, 2026 Western AI services now sit inside Iran's cyber and military workflow. The FT says Iranian military and intelligence-linked operators use ChatGPT and Gemini to su
AI News

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai