Endor Labs said on October 2 that OpenAI's GPT-6.1 Sol, run in Codex, passed its security tests on 34.1% of 200 tasks. GPT-6 Astra passed 34.6%. So both failed on about two tasks in three. GPT-6 Sol, out a week earlier, passed 25.1%. For code that only had to work, the scores were 77.7% for 6.1 Sol and 82.1% for Astra. Endor put 6.1 Sol's median task time at 8 minutes, against about 11.5 for the other two.
Each task uses code from a past security fix. The model is not told this. It is only asked to follow good security practice. A security pass needs working code first, then hidden tests from the original fix. Of the tasks 6.1 Sol got working, 43.9% also passed security. Astra solved 7 tasks securely that 6.1 Sol missed, but 6.1 Sol's code worked on 5 of them. Endor has no final Azure bill yet. It estimates $70 to $85 from token counts. The same tasks cost $104 on GPT-6 Sol and $468 on Astra.
A one-task gap matters less than the failure rate both models share. Endor says Astra's few extra working-code points come at roughly five times the token price. Buyers can ask whether the test resembles their own code, and who checks model output for security. These are one vendor's results.