Skip to content
Security & Trust

Microsoft's AI bug hunters say finding flaws is no longer the only limit

Microsoft's FORGE Lab reports 140 Windows CVEs from AI-assisted research, and says validation and fixes must keep pace.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Microsoft's AI bug hunters say finding flaws is no longer the only limit
In brief
  • Microsoft's FORGE Lab says AI agents find flaws at meaningful volume, but the value comes only if teams can validate and fix them as fast.
  • Microsoft's cost figures cover only the automated checking step, not human review or patching, so the full cost of fixing remains unstated.
  • Before buying AI-assisted testing, ask how many reports your team can validate and fix each week.

The scarce thing is not always the bug

Finding software flaws has long been seen as the hard part of security work. Microsoft's FORGE Lab now says finding is no longer the only limit. Its AI agents find flaws at meaningful volume. That work pays off only if teams can confirm and fix the flaws at the same pace.

Microsoft does not say finding got easy. It says that uncovering a hard flaw still shows what an AI agent can do. The scarce resource, it adds, can be a working build, a reproducible trigger or an engineer's time.

These figures come from Microsoft's own blog post, and WebPulse has not independently verified them. Some reports have public CVEs and maintainer acknowledgement. The lesson reaches past Microsoft: a found flaw is only the first step.

140
Windows CVEs from FORGE-assisted work
Source: Microsoft Security, FORGE Lab (October 7, 2026)

A CVE is a public ID for a known security flaw. Microsoft says FORGE helped find Windows flaws that received 140 CVEs from May to September 2026. Fixes for 52 of them were addressed in the September 2026 security release.

How the system works

FORGE runs a tool it calls MDASH, short for multi-model agentic scanning harness. It mixes several AI models, specialised code auditors and code-analysis tools. Work can be routed to the part that suits each task, though Microsoft says good routing is a hypothesis to test, not a given.

A scan produces candidate reports. Some are wrong, repeated or unreachable, and the team must sort them. So FORGE builds project-specific "provers". These programs try to turn a suspicion into a crash or trigger that can be repeated. A report with a working trigger gives engineers something to act on. A report without one gives them homework.

Microsoft also found that adding more auditors can raise the number of candidates without raising the number of confirmed findings. If reports arrive faster than its security response centre can resolve them, the queue grows. One internal project mapped code structure to spot repeats. That cut about 45% of duplicate findings across repeat scans.

About 45%
Duplicate findings removed by one internal project
Source: Microsoft Security, FORGE Lab (October 7, 2026)

What the machine step costs

The cost figures come from Linux kernel work. MDASH flagged thousands of suspicious spots. Microsoft's validation agents backed 627 findings with supporting evidence.

Each of 182 confirmed crashes got a working demonstration. The average was $3.61 in AI usage fees and 21.5 minutes. In six chosen cases, the team tested whether a bug could let a local user gain higher privileges. That averaged $8.56 and 25.4 minutes.

$3.61
Average model cost to build a crash proof of concept (182 Linux findings)
Source: Microsoft Security, FORGE Lab (October 7, 2026)

The limits matter. The runs used GPT-5.5 without extensive tuning. The averages cover successful cases only. They leave out early screening, failed candidates, human investigation and patch preparation. So Microsoft gives no cost for human review or patching.

Within those limits, Microsoft concludes that automated validation can give useful results at practical cost and speed, even at kernel scale.

Our inference is that people are the likely costly part of the rest. Microsoft says human attention should go to judging impact and reviewing patches. It gives no figure for that work.

Fixing depends on other people's workflows

Over three months, the team sent 155 internally validated reports to 23 open-source projects. Maintainers had acknowledged or accepted 93 of them when Microsoft wrote the post. Those 93 span 14 projects or project families, including curl, Node.js, SQLite and the Linux kernel.

93 of 155
Reports with maintainer acknowledgement or acceptance
Source: Microsoft Security, FORGE Lab (October 7, 2026)

Microsoft says the reports are at different stages, so few are public yet. Public examples include two curl flaws, CVE-2026-9545 and CVE-2026-13608. They also include a Node.js flaw, CVE-2026-56848, and a Linux kernel flaw, CVE-2026-64563.

Separately, Microsoft says each project works differently. Some need a patch and some do not. Some prefer private disclosure and some use public channels. Maintainers sometimes disagree on whether a flaw crosses a security boundary. Microsoft's view is that research teams should absorb this complexity, not hand it to maintainers.

What leaders should ask

This is our view, not Microsoft's: when AI makes discovery cheaper, your slowest stage sets your real speed. A scanner that finds more bugs helps only if the later stages keep up.

Before buying AI-assisted testing, put these questions to your teams.

First, how many reports can we validate and fix each week today? Second, who owns each handoff from finding to proof to patch to release? Third, when a report fails, does the tool learn why, or does a ticket just close? Fourth, how many duplicates reach a human reviewer? Fifth, if a vendor promises more findings, what happens to our review queue?

A found flaw is a debt you have acknowledged, not one you have paid.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Microsoft Security.

Share this insight