Microsoft announced that its MDASH system has outperformed competitors Claude Mythos and GPT-5.6 Sol in a recent cybersecurity test, achieving the top position on the CyberGym public benchmark for AI-driven vulnerability discovery. The success of MDASH is attributed to its multi-model approach, which utilizes specialized AI agents that collaborate in an end-to-end pipeline to identify and validate security flaws. This development highlights the growing trend of evaluating frontier AI models for addressing cybersecurity challenges, such as automated vulnerability hunting.
MDASH: MDASH is Microsoft’s codenamed multi-model agentic scanning harness, an AI-driven system that uses specialized agents to discover, validate, and help remediate software vulnerabilities across codebases like Windows. The system is currently in private preview and was highlighted in the news for topping industry benchmarks against other frontier AI models in cybersecurity tasks.
GPT-5.6: GPT-5.6 is OpenAI’s family of next-generation language models, including the flagship Sol variant designed for complex reasoning, coding, and cybersecurity applications. The news references GPT-5.6 Sol as a competing model that was surpassed by Microsoft’s MDASH in a specialized cybersecurity test.
Microsoft: Microsoft is a major technology company developing software, cloud services, and AI systems, including security tools for enterprise and consumer use. In this news, Microsoft announced that its MDASH system outperformed competing AI models in a cybersecurity benchmark test for vulnerability discovery.
Claude Mythos: Claude Mythos is a series of advanced large language models developed by Anthropic, with variants focused on high-capability tasks including cybersecurity research and vulnerability analysis. In the context of this news, Claude Mythos served as one of the competing models that MDASH outperformed in Microsoft’s cybersecurity benchmark evaluation.
System Design: MDASH employs a multi-model approach with specialized AI agents working together in an end-to-end pipeline for identifying and validating security flaws.
Industry Context: Frontier AI models from multiple companies are increasingly being evaluated and applied to cybersecurity challenges such as automated vulnerability hunting.
Benchmark Performance: Microsoft’s MDASH system achieved the top position on the CyberGym public benchmark for AI-driven vulnerability discovery.
