Skip to content
Brief The AI-First Web ·

On Endor Labs' test, GPT-6.1 Sol and Astra both fail most security checks

Endor Labs says OpenAI's new model nearly ties GPT-6 Astra, but both fall short on about two tasks in three.

In brief
  • Endor Labs found GPT-6.1 Sol passed its security tests on 34.1% of tasks and Astra on 34.6%. Both failed about two in three.
  • Endor has no final bill for the run. Its cost figure is an estimate from token counts.

Endor Labs said on October 2 that OpenAI's GPT-6.1 Sol, run in Codex, passed its security tests on 34.1% of 200 tasks. GPT-6 Astra passed 34.6%. So both failed on about two tasks in three. GPT-6 Sol, out a week earlier, passed 25.1%. For code that only had to work, the scores were 77.7% for 6.1 Sol and 82.1% for Astra. Endor put 6.1 Sol's median task time at 8 minutes, against about 11.5 for the other two.

Each task uses code from a past security fix. The model is not told this. It is only asked to follow good security practice. A security pass needs working code first, then hidden tests from the original fix. Of the tasks 6.1 Sol got working, 43.9% also passed security. Astra solved 7 tasks securely that 6.1 Sol missed, but 6.1 Sol's code worked on 5 of them. Endor has no final Azure bill yet. It estimates $70 to $85 from token counts. The same tasks cost $104 on GPT-6 Sol and $468 on Astra.

A one-task gap matters less than the failure rate both models share. Endor says Astra's few extra working-code points come at roughly five times the token price. Buyers can ask whether the test resembles their own code, and who checks model output for security. These are one vendor's results.

A WebPulse Brief: a short report of an important event, written by the WebPulse Newsroom with AI assistance and checked against the reporting below. When there is more to explain, we follow up with a full story. How we use AI.

Reporting: Endor Labs, Endor Labs.