Microsoft MDASH Multi-Agent AI System Scores 88.45% on CyberGym Security Benchmark
Microsoft has developed MDASH (Multi-Model Agentic Scanning Harness), an AI-driven security pipeline that coordinates multiple models and agents to automate vulnerability discovery. The system achieved an 88.45% score on the public CyberGym leaderboard and identified 16 CVEs during a Patch Tuesday evaluation. In a controlled test with 21 planted vulnerabilities, MDASH detected all of them with zero false positives, while also achieving 96% and 100% recall against five years of historical Microsoft Security Response Center cases for clfs.sys and tcpip.sys respectively. Unlike single-model approaches, MDASH chains together multiple steps including code analysis, hypothesis generation, exploitation testing, and result validation. Security experts note that enterprises adopting such agentic systems must also establish clear governance over model access, output logging, and human review requirements.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in