- Endor Labs found GPT-6 Sol on Codex cost 78% less than GPT-6 Astra on 200 coding tasks. Its patches passed security tests 25.1% of the time, against Astra's 34.6%.
- This is one vendor's single-run benchmark. Our argument, not Endor's finding, is that if cheaper models mean more patches, review and testing become the limit.
- Ask who reviews AI patches, whether your tests cover security cases, and what the pass rate is on your own code.
An old economic idea says that when something gets cheaper, people use more of it. If that holds for AI-written code, the bottleneck moves from writing patches to checking them. That is our argument, not a finding from the benchmark below. The benchmark does show how thin the secure share of one model's patches was on this task set.
What Endor Labs measured
Endor Labs, a software supply chain security vendor, tested Codex running OpenAI's GPT-6 Sol. OpenAI released the model on September 22, 2026. The test used 200 real coding tasks from Endor's Agent Security League.
Each task contains code that was once part of a security fix. The agent is not told a vulnerability is involved. It receives only a general instruction to code securely.
The result: 72.1% of patches worked (Endor calls this FuncPass). Only 25.1% both worked and passed hidden security tests (SecPass). That ranks sixth among the agent and model pairings Endor has tested.
Cheaper than Astra, but 9.5 points behind on security
The full run cost $104 on Azure. Endor says that is the cheapest Codex run it has measured. GPT-6 Astra, a higher-tier OpenAI model, cost $468 on the same tasks.
Tokens are the chunks of text a model reads and writes, and they are what vendors bill for. Sol processed 15% more of them than Astra. Its per-token prices are 80% below Astra's, which more than covers the extra volume.
The security gap to Astra remained. Astra scored 34.6% on SecPass, 9.5 points higher. GPT-6 Sol beat its predecessor, GPT-5.6 Sol, by 5 points. Endor calls that improvement modest but real.
Here is how we read it. On this run, the model bill was the smaller worry. The larger question was how much of the output could be trusted, and for that the tests did the work. If cheaper models lead teams to produce more patches, review capacity and test quality become the limit. Endor did not measure that effect.
How security tests catch what working code misses
A patch can pass every functional test and still be unsafe. Endor's Plone example shows how. Plone is an open-source content management system. The task involved CVE-2015-7316, in a function called isURLInPortal that decides whether a link points inside the site.
Many agent and model pairings got the link logic right. Only GPT-6 Sol also passed the hidden tests that reject script-injection payloads, such as script tags and javascript: links. The original fix blocked four known attack strings.
The agent took a different route. Its patch peels back URL encoding up to 10 times, so a payload disguised by double or triple encoding still shows up before the check. It then rejects control characters, angle brackets, quotes and backslashes. Endor calls this approach arguably more robust than the blocklist.
The lesson is that working code and safe code are separate questions. Only the hidden security tests separated the winner from the rest. Teams relying only on functional tests would likely have accepted the other patches.
Caveats on the evidence
This is one vendor's benchmark. Each agent ran each task once, using one tool (the Codex command-line harness). The scores describe this task set, not every codebase.
Endor also checked for cheating, meaning the model recovering a known fix from git history, the web or memory. It inspected 10 flagged GPT-6 Sol instances and confirmed none. Confirmed cheating is removed from scores.
Endor said OpenAI has since released GPT-6.1 Sol, and that results are not yet in.
Questions to put to your team
First, ask who reviews AI-generated patches, and how many each reviewer can check per week. If cheaper models mean more patches, that volume may grow.
Second, ask whether your tests include security cases, or only checks that the feature works. In Endor's Plone example, the security tests were what set one patch apart.
Third, ask whether you have measured the pass rate on your own code. A public score of 25.1% says little about your repositories.
Price is one cost of AI-written code. Checking it is another, and this benchmark shows why that second cost deserves its own plan.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Endor Labs.





