Skip to content
Security & Trust

AWS shows how to make AI check scanner findings before engineers see them

A three-layer design treats engineer trust as the scarce resource and verifies AI claims against the code first.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
AWS shows how to make AI check scanner findings before engineers see them
In brief
  • AWS Security published a three-filter design that checks scanner findings for plausibility, code structure and deployment context before engineers review them.
  • The lesson is that engineer trust is the scarce resource, so AI-generated claims should be tested against the code before a person sees them.
  • Leaders should ask which filters their own triage has, how independent their scanners are, and who reviews the top-priority findings.

Security teams run on trust, and trust runs out. If engineers learn that alerts are often wrong, they learn to doubt the tool. Real alerts then suffer along with the false ones.

AWS Security described a way to protect that trust in a post on October 7, 2026. It is Part 1 of a series on an "AI vulnerability harness". The design turns a flood of scanner alerts into a short list. Each item on the list comes with evidence.

Three filters, each asking a different question

AWS argues that three narrow filters beat one broad one. Each filter looks for a different kind of mistake.

The first filter asks whether a finding is plausible. Several scanners examine the same code on their own. If two or more report the same issue, confidence rises. If only one does, the finding gets a closer look. AWS calls this multi-scanner agreement. It works with tools a team already owns.

2 or more
Scanners that agree to raise confidence in a finding
Source: AWS Security Blog (October 7, 2026)

The second filter asks whether the finding is real in the code. AI scanners often describe an attack route. Input enters one function, moves through another and ends somewhere harmful. AWS says these descriptions can be wrong.

So the pipeline tests them. Is the named file there? Is the named function there? Can data really follow that route? If any answer is no, the finding is dropped.

The test relies on the code's abstract syntax tree. That is a parsed outline of how the program is built. The check runs against the parsed code structure rather than relying on the model's own say-so.

The third filter adds deployment context. The pipeline reads infrastructure-as-code files, such as CloudFormation or Terraform templates, next to the application code. It looks for protections such as AWS WAF rules, network isolation, authentication and input validation.

AWS does not treat a protection as simply on or off. A firewall rule that blocks the exact attack technique counts for more than one aimed at a different flaw.

3
Filtering layers: plausibility, structure, deployment context
Source: AWS Security Blog (October 7, 2026)

The idea: let the model argue, let the code check

Think of a newsroom. A reporter can be brilliant and still get a name wrong. Good editors do not rely on the reporter's confidence. They check the claim before it prints.

This pipeline gives AI the same treatment. The model brings reasoning that pattern-matching scanners lack. It can follow data across files and weigh how the system is deployed. But its output is a claim to be verified.

AWS names the price of skipping that step. When wrong findings land in front of people, reviewers lose faith in the tool. A tool nobody believes protects less.

The third filter also changes how to rank work. AWS notes that a top-severity finding behind strong controls may matter less than a medium finding exposed directly to the internet. A severity label alone does not set priority. Evidence of exposure does.

What the harness does not do

AWS is direct about the limits. They matter to anyone planning to fund this approach.

The result is a ranking of findings that look exploitable. None is a confirmed attack. AWS says top-priority (P0) findings still deserve an engineer's review before they drive code changes.

The harness also cannot find a class of flaw that none of your scanners look for. And reading infrastructure templates shows what was defined, not what is live. A WAF rule switched off last week would not show up. Checking live state needs a separate integration.

One caveat on the evidence. The post describes a design. It gives no figures for how many findings each filter removes. It calls the noise reduction "substantial" without measuring it. Read it as architecture guidance, not a benchmark.

The confidence formula and the verification gates sit in a companion post on configuring the model's instructions. That post was not part of this source.

Questions to put to your security team

AWS suggests building in order of payoff. Start with an index of entry points and sinks. Next, triage the output of tools like Semgrep or CodeQL. Then add AI-driven discovery, then infrastructure context. Live testing against a pre-production target comes last.

5
Suggested build steps, from indexing to live verification
Source: AWS Security Blog (October 7, 2026)

You do not need AWS's exact design to use its logic. Five questions follow from it.

First, does anything check an AI-generated finding against the code before an engineer reads it? Second, how independent are our scanners? Do they share the same blind spots? Third, does our ranking use deployment controls, or only severity scores?

Fourth, how do we confirm a protective control is still switched on? Fifth, who signs off on the top-priority items?

AWS puts the point in one line: "Configuration is what turns capability into judgment." The same holds for the people reading the results. A tool that earns engineers' trust gets acted on. One that spends it does not.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: AWS Security.

Share this insight