Skip to content
Security & Trust

AWS says its agent found, proved and patched flaws in 89% of C/C++ test tasks

The score is self-reported and covers one bug class. The design points to where security work may be moving: toward evidence.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
AWS says its agent found, proved and patched flaws in 89% of C/C++ test tasks
In brief
  • AWS reports its Continuum agent passed 819 of 920 tasks (89.0%) on CyberGym-E2E, a benchmark limited to memory-safety flaws in C and C++. AWS ran the test itself.
  • AWS describes a design in which each finding carries a working proof and a tested patch. If that holds in practice, human review becomes the bottleneck.
  • Ask vendors what their results cover, who verified them, and how findings are ranked in your own environment.

A security alert is a claim. A crash input plus a patch that passes the project's tests is evidence. AWS's latest result shows how far apart those two are, and why that gap shapes how security teams spend their time.

What AWS reported

On October 5, AWS Security published results for AWS Continuum for code vulnerabilities. The test was a public benchmark called CyberGym-E2E. AWS says Continuum cleared 819 of the benchmark's 920 tasks inside the 90-minute cap.

89.0%
Pass rate on the benchmark's main measure (S3)
Source: AWS Security blog (October 5, 2026)

A pass on this measure has three parts. The agent submits an input that crashes the program. Its patch stops that crash. The project's own tests still pass.

The pass does not mean the agent fixed the exact bug the benchmark had in mind. A codebase can hold several real flaws, so the agent may have repaired a different one. A separate, diagnostic stage checks for the benchmark's chosen flaw. The source gives no figure for it.

AWS says the previous public high was 65.9%, a gap of 23.1 percentage points. With no time cap, AWS reports the pass rate reached 93.7%.

AWS ran the evaluation itself. It says outside network access was blocked. It also says a review of the run records showed the agent analysed the code instead of looking up known public fixes. The post does not mention independent replication.

How the test works

Many benchmarks score a single step. Some check whether a tool can spot suspect code. Others start from a flaw someone has already found and ask for a patch. CyberGym-E2E scores the whole chain, from discovery to repair.

In each task, the agent works in a sandbox holding an older, vulnerable copy of a real open-source project. It is given no description of the flaw, no proof of concept and no crash log. Within 90 minutes it must hand back a triggering input and a patch.

The benchmark draws on historical OSS-Fuzz vulnerabilities. OSS-Fuzz is a public effort that finds bugs by feeding programs random inputs. The median project in the set has more than 600,000 lines of code.

920 across 139 projects
Tasks in the benchmark
Source: AWS Security blog (October 5, 2026)

How Continuum is built

AWS describes a multi-agent system with three phases: discovery, validation and remediation. Discovery proposes candidate flaws and records the source evidence for each. Validation builds a proof of concept and runs it against the vulnerable program. Remediation traces the root cause, writes a patch and checks it against the same proof.

The key design choice is that evidence travels with the finding. The patch is judged against the crash that justified it. That matters because a reviewer can check a claim instead of debating it.

What the score does not cover

AWS is direct about the limits. The benchmark covers memory-safety flaws in C and C++ only. Sanitizers, which are tools that catch memory errors as a program runs, give objective proof of a crash.

AWS concedes that other languages and other bug types sit outside the test. It says that includes many of the most common and impactful bug classes AWS sees in production.

The benchmark also gives no deployment context. AWS notes that how urgent a flaw is depends on how the service is exposed, its network paths, its permissions and its settings. AWS says Continuum can use that context for customers. CyberGym-E2E does not test it.

So 89.0% is a result for one kind of bug in one kind of code. It does not say how the system performs on your applications.

The lesson: evidence is the new unit of work

AWS frames the problem as volume. Potential vulnerabilities now arrive faster than existing processes can investigate, reproduce and repair them.

This shows where the effort could move. AWS describes findings that carry a proof and a tested patch. If that holds in customer environments, review becomes the bottleneck. The scarce skill is then judgment: is this fix right, is this flaw reachable, does it matter here?

Someone must own that judgment. A reviewer who once chased a vague report would check a crash and a patch instead. That may be better work, but it is still work.

It also changes how to read vendor claims. A single success rate hides what was measured. The useful questions are about scope, verification and context.

What to ask your team and vendors

First, ask what share of your code is C or C++ with memory-safety exposure. That is where this benchmark applies. Second, ask any vendor quoting a benchmark what it excludes, and whether anyone outside the company has reproduced the result.

Third, ask who approves an AI-written patch before it ships, and what evidence they see. Fourth, ask how findings are ranked once deployment context is included, since the benchmark does not test that step.

A fix with proof is easier to trust. The harder test is whether your team knows which fixes matter first.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: AWS Security.

Share this insight