Cost per task, Sonnet 5.5 (Artificial Analysis benchmark, max effort): $7.60 (Source: Artificial Analysis, Claude Sonnet 5.5 evaluation (September 28, 2026))
The rate card is not the bill
Most AI budgets start from a rate card: dollars per million tokens. Artificial Analysis' evaluation of Claude Sonnet 5.5 shows why that figure can mislead. Anthropic priced Sonnet 5.5 the same as Sonnet 5, at $2 per million input tokens and $10 per million output tokens. In the max-effort results on Artificial Analysis' Intelligence Index, however, the benchmark cost per task comes to $7.60, about 50% above Sonnet 5's.
The price per token did not rise. The number of tokens each benchmark task consumed did. For anyone approving spend on AI agents, the lesson is that the number that matters for budgeting is the cost of a completed task, and the rate card does not show it. The invoice bills tokens, so rising token volume is how a change like this would appear.
Where the extra cost comes from
At max effort, Artificial Analysis measured about 193,000 output tokens per Intelligence Index task. The firm says this is the highest token use it has measured. Its volume exceeds Opus 5.5's and Sonnet 5's at max effort by around 60%, and GPT-6 Astra's by a factor of about seven.
The extra work bought capability. Artificial Analysis credits max-effort Sonnet 5.5 with an 18-point improvement on Sonnet 5, placing it #2 on the Intelligence Index, with only Opus 5.5 at max effort above it. Terminal-Bench 4.0 tests agentic terminal use; there Sonnet 5.5 posted 64%, slightly above the 60% listed for Opus 5.5 and for GPT-6 Astra at xhigh. Artificial Analysis describes that as a 50-point increase over Sonnet 5 at max. On three knowledge-work measures (AA-Briefcase, GDPval-AA and AutomationBench-AA), Sonnet 5.5 reached parity with Opus 5.5, though Artificial Analysis notes it used significantly more tokens to get there.
Effort is now a purchasing decision
Sonnet 5.5 offers five effort settings, from low to max, and Artificial Analysis ran its Intelligence Index at all five. At this pricing, it reports, Sonnet 5.5 sits off the Intelligence versus Cost per Task Pareto frontier. At lower efforts, GPT-6 Astra or Sol configurations deliver equivalent performance for less. Of the five settings, high effort makes the strongest case on cost against intelligence: it trails GPT-6 Sol by a very small margin while costing effectively the same per task.
So the choice is not only which model to buy. Someone must also decide how hard the model works on each job, and the source shows those settings landing at different points on the cost curve. Does anyone in your organisation own that setting for each production workflow?
What the score does not settle
Two limits matter before anyone acts on the ranking. First, Sonnet 5.5 trails Opus 5.5 on factual knowledge. On AA-Omniscience it scores 54% for factual accuracy against 66%, though its hallucination rate is lower (47% against 59%). It also sits about six points below Opus on Humanity's Last Exam and SciCode. Second, the tests ran on a pre-release deployment that Anthropic found had a bug affecting requests that use structured outputs. Anthropic expects minimal change or slightly understated performance, per Artificial Analysis, which plans to re-run the relevant evaluations. Treat these figures as a first reading.
These are also benchmark tasks, not your workloads. Your cost per task will depend on what your agents are asked to do.
Questions for your team
First, ask whether AI spend is tracked per completed task as well as per token. The per-token rate alone would not have shown this change; total token volume would. Second, ask who sets the effort level on each production workflow and whether anyone has compared cost at two settings on the same job. Third, ask whether contracts and forecasts assume that an unchanged rate card means an unchanged bill. Fourth, ask whether a re-test is scheduled once Artificial Analysis publishes its planned re-run.
A taxi with an unchanged per-mile rate still costs more if the route gets longer. With AI agents, the meter is the number of tokens a task consumes, and the budget owner should be watching it.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Artificial Analysis.





