- Snyk tested its live-attack tool against a code-reading AI on one vulnerable app. Its tool confirmed 10 of 15 exploit chains. Snyk ran the test.
- Two flaws that looked separate in a code review combined into admin account takeover. Severity lists do not show that.
- Ask your vendors whether findings are proven against the running system, and whether chains are tested. Treat vendor benchmarks as one data point.
A list of flaws is not a route into your systems
Most security reports are inventories. They list what is wrong and rank each item by severity. Attackers do not work from inventories. They look for a path, and a path can be built from items that each look minor.
That is the idea at the center of a test Snyk published on October 7. Snyk argues that attackers do not read your repository. They hit your URL and combine whatever they find. This is Snyk's view, and it is a vendor's view. The test behind it is worth understanding.
What Snyk tested
Snyk pointed two tools at one target. The target was TaintedPort, a deliberately vulnerable web app that Snyk builds and maintains. Snyk graded each tool against a fixed checklist of 57 documented flaws and 15 documented chains. An exploit chain is a set of flaws that work together to reach a goal that none reaches alone.
The first tool was Snyk's own Evo Continuous Offensive Security (COS). It attacks the running application through a live URL. In this run it also had the source code. The second was Claude Security running the Mythos model. It read the source code only.
These are different kinds of tools. Snyk says so itself and likens the split to the old one between dynamic and static testing. One reads the code. The other attacks the running system.
How two flaws became an admin takeover
Snyk gives one worked example, labeled CHAIN-004. Every tool in the test found two flaws. The first was a server-side request forgery (SSRF), which lets an attacker make the application fetch data on the attacker's behalf. The second was a hardcoded signing secret for JWTs, the tokens that prove who a user is.
Alone, each is a finding with its own severity score. Together, Snyk says, they are an account takeover. Evo COS went further than the other tools. It abused the request-forgery flaw to pull the secret out of the live app. It then signed its own administrator token with that secret, which gave it control of the admin area. Snyk says it attached a runnable proof of concept.
Snyk explains why code reading struggles here. To confirm a chain, a tool must walk through each step on the live app and check that the next door actually opens. A model reading source has to guess whether a flaw really fires, and in Snyk's words it "misses the chain."
The numbers, and who produced them
Snyk scored each tool on a severity-weighted scale. Low counted 1 point, Medium 3, High 9 and Critical 27. Confirmed chains scored at their own severity on top of their member findings. The total was 938 points: 587 from vulnerabilities and 351 from chains.
Snyk reports that Evo COS found the most vulnerabilities. It reports precision of 96.2% against 90.2% for Claude Security. Precision means how many reported findings were real. On critical-severity flaws, Snyk's count favored Claude Security by a single finding, 10 to 9.
Snyk also says Claude Security caught several flaws Evo COS missed. These were mostly source-visible logic and cryptographic flaws. Snyk concedes neither approach is complete on its own.
Read the limits before the result
This is a vendor comparing its own product on its own test app. Snyk says Evo COS was not tuned against TaintedPort, and the app is public for anyone to reproduce. Even so, the test is one run per tool on one application. Snyk states the results are single-run outputs, not medians.
The tools also did different jobs. A live attack is a more intensive process than one pass over code. Snyk acknowledges this. The comparison shows what live testing can prove. It does not rank the underlying models.
What to ask your security team
The useful lesson is about evidence, not about one vendor. A severity count tells you how much is wrong. It does not tell you what an attacker can reach. Put these questions to your team and your vendors.
First, which of our findings has anyone proven against the running system, and which are inferred from code? Second, do our reports show chains, or only individual flaws? Third, when two medium-rated issues share an application, who checks whether they combine? Fourth, who validates a finding, and is it the same system that generated it?
Snyk's own point on that last question is fair: the system that generates a finding should not be the one that grades it. Ask any tool you buy to show its answer.
In CHAIN-004, two findings that every tool reported became an admin takeover only when a tool connected them. Your reporting should be able to show that connection before an attacker does.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Snyk.





