
Security
MDASH stuffed with MAI-Cyber-1-Flash and a side of GPT-5.4
AI agents can break through security, but they are also the solution to defending against an increasingly dangerous ecosystem of threats. On Monday, Microsoft announced a new security model that it says helped outperform several rival AI systems on a vulnerability benchmark while cutting costs by about half.
Unsurprisingly, at Redmond’s security event on Monday, execs touted the tech giant’s AI security prowess and introduced a new agentic security system called Project Perception, and also unveiled its first security-specialized model, MAI-Cyber-1-Flash, designed for software vulnerability analysis.
Microsoft packed MAI-Cyber-1-Flash, based on Microsoft AI (MAI)’s internally developed MAI-Thinking-1 reasoning model, inside its MDASH bug-hunting harness. Its execs claim the duo – with a GPT-5.4 boost – outperforms Anthropic’s bug-hunting machine Mythos and OpenAI’s powerful standalone models, and costs about half the price of other leading commercial models.
CyberGym’s benchmarking found that MAI-Cyber-1-Flash, combined with GPT-5.4, both stuffed inside the MDASH harness, achieved a 95.95 percent success rate. For comparison, OpenAI’s GPT-5.5 Cyber scored 85.6 percent and its GPT-5.6 Sol scored 83.6 percent, while Anthropic’s Mythos 5 successfully handled real-world vulnerabilities 83.8 percent of the time. Google’s Gemini 3.5 Flash Cyber in CodeMender achieved an 83.2 percent success rate.
“This is really quite a remarkable result,” Mustafa Suleyman, CEO of Microsoft AI, said during the Monday event.
Within MDASH, MAI-Cyber-1-Flash handles up to 90 percent of all queries, detecting and patching the vulnerabilities while also confirming the fixes worked, and hands the remaining 10 percent of tasks off to the larger GPT-5.4, Suleyman explained.
“GPT 5.4, which is obviously a larger model, about 10X larger, solves those [queries],” he said. “As the models hand off between each other, they are not just able to deliver better performance than all of the other models combined, they do so at 50 percent of the cost.”
In addition to the multi-model bug hunting system, Microsoft announced Project Perception, which coordinates three types of agents: red team agents that find and simulate attack paths, blue team agents that investigate and determine risk, and green team agents that remediate the issues.
“We need to make sure that the defenders can defend at the scale and the speed of the attackers,” Hayete Gallot, executive vice president of Microsoft Security, said. “You need a new cyber stack. So we built it. This is what we call Perception.”
Aside from the new security products, Redmond introduced a new AI security research arm called Microsoft Security FORGE (Frontier Offensive Research and Generative Exploration) Labs, led by Microsoft VP of Security Research Taesoo Kim, and an AI red team alliance.
The latter, called the External Red Team Alliance (EXTRA), aims to expand AI safety research through a two-part initiative.
First, Redmond’s own AI red team provided “unrestricted gifts ” to 18 university labs across six continents to support AI safety research, Microsoft data cowboy and AI red team lead Ram Shankar Siva Kumar said in a blog.
MORE CONTEXT
“The funding is unrestricted because the objective is not to direct research outcomes toward product requirements or predefined deliverables,” he wrote. “Some universities are examining the cybersecurity implications of AI systems themselves – including how models can be attacked, manipulated, or abused in operational environments. Other labs are exploring the inverse problem: how AI systems can assist defenders and improve cyber operations.”
The second EXTRA component will build a distributed network of specialists to participate in red teaming across very specific areas. “That includes researchers, practitioners, and regional experts who understand specific attack classes, languages, cultural contexts, or technical domains that internal teams may not fully cover alone,” he added. ®