- Endor Labs found OpenAI's GPT-6.1 Sol, run on Codex, scored 34.1% on secure code, against 34.6% for the pricier GPT-6 Astra.
- Both scores are low. Roughly two in three benchmark tasks were not solved securely, so a cheaper model does not remove the need for review.
- Test any model on your own code before choosing it, and track security results separately from price.
For years, buyers have assumed that a higher price buys a safer product. One new benchmark result from Endor Labs strains that assumption for AI coding tools. It also shows how much security work remains undone, whichever model you pick.
What Endor Labs found
Endor Labs is a software supply chain security vendor. Its researchers ran OpenAI's GPT-6.1 Sol through the Codex coding harness and scored it in their Agent Security League. They used the same coding tasks they had given to GPT-6 Sol and GPT-6 Astra.
On security, GPT-6.1 Sol scored 34.1% SecPass. Astra scored 34.6%. That is a one-task difference. GPT-6 Sol, released a week earlier, scored 25.1%. Endor Labs places GPT-6.1 Sol third on its current leaderboard, directly behind Astra.
The price gap
OpenAI describes GPT-6.1 Sol as near-Astra intelligence at one-fifth of Astra's price. Endor Labs puts the difference at roughly five-fold in token price. Tokens are the small chunks of text an AI model reads and writes, and they are how usage is billed.
Billing is a separate question. Endor Labs says its invoice for the GPT-6.1 Sol run is not final. From token counts and list prices, it estimates $70–85 for the full run. For comparison, the earlier invoices on these tasks were $468 for Astra and $104 for GPT-6 Sol.
This shows one cheaper model matching a pricier one on security. Price did still track security inside the GPT-6 family: the cheaper GPT-6 Sol scored 25.1%, well below Astra's 34.6%. The link between tier and security simply did not hold for this newest model. That is one data point, not a rule.
How the test works
The researchers gave each model real coding tasks inside a code repository. Each model's fix was applied as a patch and then tested. FuncPass measures whether the code works. SecPass measures whether the fix is also secure.
Benchmarks can be gamed, so Endor Labs also looks for signs of cheating. Five signals flagged 13 instances, including patches that looked too similar to known ones and signs of memorised answers. Two independent AI review rounds and a tiebreaker examined each case. All 13 were cleared.
That makes three Codex runs in a row with no confirmed cheating. It is also why Endor Labs reports identical raw and adjusted scores.
Speed matters too
GPT-6.1 Sol had a median of 8 minutes per task. GPT-6 Sol and Astra took about 11.5. The four-worker run took 13 hours, against 19 to 20 hours for the other two. Faster runs mean less time waiting on automated coding agents.
The number behind the headline
The near-tie is real, but look at the score. Astra's 34.6% SecPass means roughly 65% of tasks were not solved securely. GPT-6.1 Sol's 34.1% leaves a similar share. Astra solved 7 tasks securely that GPT-6.1 Sol did not. GPT-6.1 Sol solved 6 that Astra did not.
On functional correctness, Astra led by 4.5 points, or eight tasks. So the models differ a little on whether code works. They barely differ on whether it is secure. Both leave most of the security work undone.
Some caveats apply. This is one vendor's benchmark, run on one harness, using Codex CLI 0.144.6. The cost figure for GPT-6.1 Sol is an estimate. Endor Labs compares only the GPT-6 Codex runs here, so these scores say nothing about every model on the market. A single set of tasks cannot tell you how a model performs on your codebase.
What leaders should ask
First, ask your engineering team which model writes your code and why. If the answer is price or brand, ask for security test results on your own repositories.
Second, ask whether security failures are tracked separately from functional failures. Code that works but is not secure passes ordinary tests and ships.
Third, ask who reviews AI-written code before release. With Astra at 34.6% and GPT-6.1 Sol at 34.1%, the review step is where much of the security risk is caught or missed.
Fourth, ask whether the team could switch models if a cheaper one tests as well. Endor Labs saw the gap close in a week, so a contract that locks in one tier may carry its own cost.
Price tells you what a model costs. Only a test on your own code tells you what it secures.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Endor Labs.





