HackMyIP
← Back to News
2026-07-28 The Hacker News

Microsoft's MAI-Cyber-1-Flash Hits 95.95% on CyberGym, Halves MDASH Cost

AI SecurityVulnerability

Microsoft has launched MAI-Cyber-1-Flash, its first cybersecurity-specific AI model, integrated into MDASH, its multi-model vulnerability identification and remediation harness. According to Microsoft, the configuration pairing MAI-Cyber-1-Flash with GPT-5.4 achieved a 95.95% score on the CyberGym Level 1 benchmark while cutting costs by 50% compared to Microsoft's previous best MDASH setup, which combined GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex. The new model is a sparse mixture-of-experts transformer with 137 billion total parameters, 5 billion active parameters, and a 256,000-token context window, fine-tuned from MAI-Code-1-Flash and built on a MAI-Thinking-1 mid-training checkpoint.

The 95.95% score belongs to the MDASH system as a whole, not to MAI-Cyber-1-Flash in isolation. Microsoft says the new model can handle up to 90% of MDASH tasks, with GPT-5.4 reserved for the hardest 10%, while the model card notes the evaluated configuration replaced roughly 80% of the harness's existing models and raised the reported CyberGym result from 88.4% to 95.95%. This routing-based design is the central technical claim: most tasks are delegated to the smaller fine-tuned model, while GPT-5.4 picks up the remainder, a pattern that mirrors the cost-saving logic behind lightweight privacy checkup tooling that offloads only the heaviest inspection to a premium engine.

Important context tempers the headline. CyberGym Level 1 is a known-vulnerability reproduction test that gives an agent a vulnerability description and the corresponding unpatched source code, then checks whether it can produce a working proof of concept. It does not measure blind vulnerability discovery or patch correctness. When checked on July 28, 2026, CyberGym's public leaderboard did not include the 95.95% result; Wiz's Atlas agent held the top spot at 90.9%, while Microsoft's previous MDASH submission from May 12 remained at 88.4%. Microsoft has not confirmed whether the new result was submitted for listing, and a separate 96.55% figure from June used a broader scoring criterion that counted any crash, including non-target vulnerabilities, making direct before-and-after comparisons unreliable.

Access is restricted to approved MDASH customers through an Azure AI Foundry private preview, and MAI-Cyber-1-Flash is not available as a standalone public model or via a general-purpose API. The announcement does not disclose token usage, call volume, latency, or task-level performance breakdowns, leaving open questions about how the gains translate beyond the benchmark. Security teams evaluating AI-assisted vulnerability workflows should harden their own perimeter first, starting with a port scanner review of any new Azure endpoints and a password checker pass on the service principals used to invoke MDASH, before trusting preview-tier AI output to triage production code.

Source: The Hacker News →

Related Tools

Check whether this kind of story affects you — free, no signup:

My IP →IP Lookup →Privacy Checkup →

Related Guides

Learn the background behind this story:

What is my IP and why it matters →IP address security →How to stop being tracked online →