Skip to content
The AI-First Web

One benchmark: 55% of failed AI patches fixed the main flaw, left another open

Artificial Analysis's Cyber Index details how AI agents fail at finding and patching flaws, and what refusals hide

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
One benchmark: 55% of failed AI patches fixed the main flaw, left another open

AI-generated image for WebPulse. About our images

Key finding

Failed CWE-Bench-AA attempts (excluding refusals and timeouts) that fixed the main flaw but left a related one open: 55% (Source: Artificial Analysis Cyber Index (September 28, 2026))

Artificial Analysis launched the Cyber Index and a partner group, the Cyber Index Alliance, on September 28. The Alliance includes Collinear AI, IBM, Nvidia and Vercel. The Index scores AI models across three steps of defensive work: spotting weaknesses in a codebase, confirming them, and fixing them while keeping the software working. The publisher positions it as a buying aid for teams selecting models for security work. The failure data inside it arguably matters more to a budget-holder than any leaderboard position, because it bears on two decisions covered below: how patches are accepted and how much code is scanned.

The fix that closes one door

CWE-Bench-AA is an audit-and-patch test built on 120 held-out tasks. Artificial Analysis reports that partial fixes were the leading failure mode. Once refusals and timeouts are set aside, 55% of the failures were partial repairs: the agent dealt with the flaw it found while a related weakness, such as a second way in, stayed exploitable. A ticket marked resolved on that basis leaves the exposure in place.

55%
Failed CWE-Bench-AA attempts (excluding refusals and timeouts) that fixed the main flaw but left a related one open
Source: Artificial Analysis Cyber Index (September 28, 2026)

Over-correction, where a patch damages legitimate behavior, is the second pattern, at roughly 24% of failed attempts. Its share of failures rises with model strength: about 40% for the four top scorers, against about 15% for the weakest performers. These are shares of failures, not rates of over-correction per attempt. The damage was typically an edge case of legitimate behavior near the fix. The leaders' failures skew more toward over-correction than the weakest models' do, so an acceptance process that trusts a patch because the model is capable has a gap.

Discovery is where the coverage runs out

DeepsecBench-AA is Artificial Analysis's version of Vercel's DeepsecBench. Models receive files that a scanner flagged and report what they can verify as real vulnerabilities, with a human-expert-verified list as the answer key. The highest-scoring model found only 41% of those issues. The findings skewed toward flaws where hostile input leads straight to harm. Flaws that require reasoning across several steps, or about business and privacy rules, were seldom reported.

Agents also often report findings outside the expert-verified list, some of them real, and Artificial Analysis notes that a security team has to triage every one. Discovery output is therefore a workload for people as well as a score.

41%
Expert-verified issues identified by the highest-scoring model on DeepsecBench-AA
Source: Artificial Analysis Cyber Index (September 28, 2026)

There is one qualification. When models did report sequence-of-events flaws, 95% of the reports were correct. GPT-6 Sol and GPT-6 Astra found them in about 30% of runs. Artificial Analysis says this may be an emerging capability. That is one benchmark's reading, not a settled trend.

Refusals, timeouts and the wrong bug

CyberGym-E2E-AA has 131 memory-safety tasks in C/C++ projects. Six models declined 98% or more of them. The refusers span GPT-6 (Astra and Sol), Claude (Fable 5.1 and Opus 5.5) and two Qwen3.8 models. Artificial Analysis says this makes frontier performance hard to assess.

For the models that did engage, discovery was the main obstacle. In 42% of their attempts, the 90-minute limit ran out before the agent produced any input that crashed the program. The publisher ties this to a practical decision: how much of a codebase an enterprise should point these tools at.

42%
CyberGym-E2E-AA attempts (non-refusing models) that hit the 90-minute limit without a crashing input
Source: Artificial Analysis Cyber Index (September 28, 2026)

Passing attempts carry their own problem: 31% patched a real crash other than the target. Models tended to stop at the first crash they could validate. Null-pointer crashes, a shallower class of issue, made up 22% of these off-target passes.

31%
Passing CyberGym-E2E-AA attempts that patched a real crash other than the target
Source: Artificial Analysis Cyber Index (September 28, 2026)

Read the scope before the scores

Models in the Index read source code from open-source projects, and none is asked to produce a working exploit. Several areas are absent today. Artificial Analysis names handling live incidents, authoring secure new code, and testing systems whose code cannot be read, such as binaries and running services, as later additions. Until those arrive, they sit outside what the Index measures. All figures here are the publisher's own, from its implementations of partner and academic benchmarks. Coverage of your environment may differ from what was tested.

Questions for your security team

1. When we accept an AI-generated patch, do we test related entry points as well as the reported one?

2. What regression coverage must a patch pass, given that over-correction makes up a larger share of failures among the strongest models in this data?

3. Do we know how each candidate model handles refusals, and are they tracked separately from failures?

4. Who triages the extra findings an agent reports, and what capacity do they have?

5. Which of our assets, such as binaries we run without source access and live services, sit outside what these benchmarks measured?

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Artificial Analysis.

Share this insight