Skip to content
The AI-First Web

Google's AI bug hunter must prove each flaw before a product team is told

PageBreak found 500+ XSS flaws in Google's own apps. A non-AI tool confirms each reported one first.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Google's AI bug hunter must prove each flaw before a product team is told
In brief
  • Google's PageBreak agent found over 500 XSS flaws in Google's own web apps. Separate non-AI validators must exploit each one before a team is told.
  • The lesson is that proof, not discovery, is now the scarce resource in AI security testing. Google reports a near-zero false positive rate.
  • Ask vendors whether a separate tool confirms each AI finding against a running system, and what their validators cannot check.

Finding a flaw is cheap. Proving it is the hard part.

AI tools can now suggest security flaws faster than any team can check them. Google's Product Security team calls much of that output "AI slop": noisy, unverified guesses that pile work onto product teams. Google's answer is a rule that the AI does not get the final word.

The rule lives inside PageBreak, an internal AI agent that tests Google's first-party web applications. It began as a pilot in November 2025 and became a full project in January 2026. The agent is not tied to one model. Google says the large majority of its use runs on Gemini, and names Gemini 3.1 Pro and Gemini 3.5 Flash as examples.

500+
Cross-site scripting flaws PageBreak found in Google first-party web apps
Source: Google Product Security blog post on PageBreak, as reported by Dark Reading (October 6, 2026)

Cross-site scripting (XSS) lets an attacker run their own code in a visitor's browser on a trusted site. Google says some of the flaws sat on sensitive domains.

How the proof step works

The agent reads code and forms a hypothesis about a flaw. It then hands that hypothesis to a validator. Google describes these validators as specialised and "non-AI-written". Each one runs a real payload against a running application to see whether the attack works.

The checks differ by flaw type. For XSS, the validator injects JavaScript and loads the page in a rendering harness to see whether the code actually runs. For SQL injection, it checks the output or the timing of database responses. For path traversal, it creates a file in a world-readable location and tests whether the app can read it.

For remote code execution, it tries methods such as a sleep delay, a written file or an outbound DNS or web request. For server-side request forgery, it looks for a request that reached an internal service.

Models can still wander down dead ends. So Google runs agents with identical seeds across many iterations, to raise the odds of finding the right attack path. Google reports a near-zero false positive rate. That is Google's own claim; the post cites no independent test.

Who carries the cost of a bad report

Think of the engineer who owns a product page. A false alarm can cost that person time and some trust in the tool. Google's design aims to protect that person. Unverified candidates are not sent to product teams.

Rickard Carlsson, CEO of Detectify, told Dark Reading that for the past couple of years, AI security tools have flagged more suspected weaknesses than teams can realistically work through. What he values in Google's design is that the agent "doesn't get the final word." A separate non-AI validator must run a real payload first.

Volume still hurts. Google says that although PageBreak aims to send only verified findings, product teams still face an unprecedented volume of reports. It plans to link PageBreak with CodeMender, its automated fix system, so teams mainly check proposed fixes. Dark Reading reports that Google did not immediately say whether all PageBreak findings have been fixed.

What the framework result shows

Google also tested apps built on its high-assurance web frameworks. These are designed to remove whole classes of web flaws by default. As of September 4, 2026, PageBreak found only 2 XSS flaws across hundreds of such apps. Google says each was confined to an internal application or a debug endpoint where hardening was incomplete.

2
XSS flaws found across hundreds of apps on Google's high-assurance frameworks
Source: Google Product Security blog post on PageBreak (data as of September 4, 2026)

This is a strong signal, with limits. Google did not publish the exact app counts, so the 500-versus-2 gap is not a controlled comparison. The lesson is still plain: the cheapest flaw to fix is the one the framework prevents.

Limits that matter to buyers

Google says its validators do not yet cover every vulnerability type or complex scenario. That risks false negatives, which are real flaws the agent misses. Google keeps unconfirmed findings internally as seeds for later scans and as a guide for building new validators. It does not send them to product teams.

Google also names advantages most firms lack. Its code sits in one huge repository, so an agent can follow a request end to end. It has live traffic data that maps web paths to source code. It has scanners that can sign in to nearly every Google web application. These results are one company's, not a benchmark.

3
Example flaws detailed in Google's companion post (all reported fixed)
Source: Dark Reading (October 6, 2026), citing Google's Bug Hunters blog post

Questions to put to your team

Ask any vendor selling AI vulnerability scanning whether a separate tool confirms each finding against a running system. Ask what share of reports arrive with proof attached. Ask which flaw types the confirming tools cannot check, and how those gaps are handled.

Ask your engineers which web flaw classes your frameworks already block by default. Fewer flaws to find means fewer reports to triage.

An AI that can accuse is easy to build. The tool worth paying for is the one that has to prove it.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Google.

Share this insight