- Ridge Security's benchmark of eight LLMs, reported by ReversingLabs, concludes that AI pen-test results depend more on the surrounding system than on the model.
- Claude Opus 4.6 reached 63% coverage at $217 per run, while Gemini 3 Flash reached 52% at about $5.42. Grok 4.5 covered the most ground, at 77%.
- Judge vendors on validated findings, cost per validated finding, how stopped tests are reported and who enforces scope, not on model rankings.
Buyers of AI security testing usually start by asking which model scores highest. Ridge Security's new benchmark points to a different question. The vendor says the software wrapped around a model shaped results more than the model's own ability. That is Ridge's conclusion. The ReversingLabs post that reported it gives no figure for how much each factor contributed.
What Ridge tested
A penetration test is an authorized, simulated attack on your own systems. Ridge ran 96 tests across eight leading language models. The targets were deliberately vulnerable environments. Each run used Ridge's own harness, the software that directs what a model does and checks its work.
Every model ran in that same harness, so the benchmark did not compare harnesses directly. The conclusion that the harness matters more is Ridge's, and executives quoted below share it.
The researchers watched whether a model could finish every stage without stalling. Each run had to move from scouting the target, to testing ideas, to adjusting attack code, to breaking in, and finally to confirming the result. Ridge says no earlier public benchmark had compared several leading models on this task.
Cost rises faster than coverage
Results varied widely by model and by price. Grok 4.5 covered the most ground, at 77%. Claude Opus 4.6 scored 63% and cost $217 per run. Gemini 3 Flash scored 52% for about $5.42. That is roughly 40 times the price for 11 more points of coverage.
Efficiency told another story. The open-source GPT-OSS-120B was the most efficient model in Ridge's results. It produced 16.9 findings per million tokens, and each run cost $2.32. Tokens are the units AI vendors bill by.
The ReversingLabs post does not explain how coverage is calculated. Treat the percentages as Ridge's own yardstick.
Ridge's researchers say frontier models deliver stronger coverage but cost substantially more per run. Smaller and open-source models give up some coverage in return for efficiency and more choice in where they run. For many organizations, the researchers say, that trade can matter more than raw coverage.
Why the system around the model matters
A pen test is a long task. The AI must track what it has tried, use tools reliably and prove every finding. Rogier Fischer, CEO of Hadrian, says those abilities belong to the whole system rather than the model.
Ridge's team divides the work in three. The model does the thinking. The harness carries out the actions. A separate check confirms each finding is real.
Fischer lists what a harness adds: limits on scope, memory, repeatable tools, validation and an audit trail. Repeatable tools behave the same way each time.
Li Zhao of Black Duck says a well-built agent on a smaller, cheaper model can often beat a frontier model. That holds when the system is tuned for offensive work. Fischer's advice on buying is short: "A leaderboard ranking is not a procurement criterion."
These views come from security-company executives quoted by ReversingLabs. They agree with Ridge. They are not independent tests.
A stopped test can look like a clean one
Ridge saw a second problem. Frontier models sometimes balked partway through a workflow, including at payload generation and other exploitation steps. This happened even when the work stayed inside agreed limits. Seemant Sehgal, CEO of BreachLock, calls that a reasonable choice by model makers, given how easily the same skills could be misused.
The buyer's risk is quiet. Kevin Surace, CEO of Token, says a halt midway can leave a gap that someone reads as a clean result. Silence can pass for safety. He advises disclosing skipped tests and finishing the work with established tools and human testers.
Gunter Ollmann, CTO of Cobalt Labs, draws a design rule. Permission should not live only inside the model. The platform should know the allowed targets, credentials, tools and actions. It should also isolate execution and keep logs.
More findings are not more security
Ollmann expects AI to make finding flaws cheaper and faster. He says that alone does not make an organization safer. The test is whether it can validate, rank and fix what turns up.
In the same research, 47% favored automation for lower-importance assets, paired with human-led testing for critical ones. The post gives no survey date or sample size.
What to ask before you buy
This is one vendor's benchmark on deliberately vulnerable targets. It does not show how any model performs on your own systems. ReversingLabs, which reported it, is also a founding member of the Agentic SOC Alliance, a group the post promotes.
The results still suggest better questions for a vendor or your security team.
Fischer's suggested yardsticks are a good start. Ask how many findings were proven real, how often the alarms were false, whether results repeat from run to run, and what each proven finding costs. Ask how the platform reports a test that stopped early. Ask who enforces scope, the platform or the model. Ask who owns each fix once findings arrive, and by when.
The lesson here is that the model is one part of what you are buying. The Agentic SOC Alliance's own design treats the model as the layer easiest to swap out. The harness, the context and the people who act on findings are harder to replace.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: ReversingLabs.





