Skip to content
The AI-First Web

New Claude and GPT models launched cheaper, yet a failed request cost $2.56

Anthropic and OpenAI released cheaper models the same day. A runaway test request shows the rate is not the bill.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
New Claude and GPT models launched cheaper, yet a failed request cost $2.56

AI-generated image for WebPulse. About our images

Key finding

Opus 5.5 price cut versus earlier Opus models: 20% (Source: Simon Willison (September 22, 2026))

The rate on the price list is not the bill

AI vendors compete on price per token. A token is a small chunk of text the model reads or writes. Your invoice is that price multiplied by how many tokens your systems consume.

The first number fell for the newest models this month. The second is set by settings your own teams choose. A test run by developer Simon Willison shows how far apart the two can be.

What changed on September 22

Willison wrote that Anthropic launched Claude Opus 5.5 that day. Roughly an hour later, OpenAI followed with GPT-6 Sol and GPT-6 Luna. These are new models launched at lower prices, not price cuts to existing ones.

Opus 5.5 costs $4 per million input tokens and $20 per million output tokens. Five earlier Opus versions, 4.5 through 5, all sat at $5 and $25. That makes the new price 20% lower. The price of reading cached text fell 60%.

At OpenAI, Willison says GPT-6 Luna halves the price of GPT-5.6 Luna, which was already cheap. Sol dropped by a similar proportion against GPT-5.6 Sol. Luna lists at $0.10 input and $0.50 output. Anthropic's current Haiku 4.5 lists at $1 and $5, ten times higher.

Two limits apply to this picture. OpenAI has already scheduled a 25% rise for GPT-5.6 in November, so the old models were being compared at temporary discount prices.

Also, the top tier has not moved. Willison notes that GPT-6 Astra and Claude Fable 5.1 still charge $10 per million input tokens and $50 per million output. In his words, the price war so far reaches the tier below them.

20%
Opus 5.5 price cut versus earlier Opus models
Source: Simon Willison (September 22, 2026)
60%
Fall in Opus 5.5 cache-read price
Source: Simon Willison (September 22, 2026)

How the meter actually runs

Three mechanisms decide what an AI workload costs.

First, input and output tokens are priced separately, and output costs more. The model's own reasoning also counts against the output limit, as the test below shows.

Second, caching. An agent that holds a long conversation keeps resending earlier text. Willison notes that in extended agent sessions, cached pricing covers more than nine in ten input tokens. A 60% cut to that price can therefore matter more than the headline rate.

Third, the reasoning setting. Models let users choose how hard they think. The top setting on Opus 5.5 is called "max".

When thinking eats the answer

Willison asked Opus 5.5 at "max" to draw an SVG image of a pelican on a bicycle. It returned no response. That had never happened before in his test.

The model kept reasoning until it used up its whole output allowance. Opus 5.5 is capped at 128,000 output tokens, and it reached that cap while still planning the drawing. He tried a second time and got the same result.

Each failed attempt cost him $2.56 and took nearly 20 minutes. The arithmetic matches: 128,000 output tokens at $20 per million is $2.56. He got nothing usable from either run.

$2.56
Cost of each failed Opus 5.5 "max" run
Source: Simon Willison (September 22, 2026)

This is one tester, one prompt and one setting. It does not show how often this happens. Willison himself says the result leaves him distrusting "max" for more serious work.

The lesson: cheaper rates do not cap exposure

This shows a gap in how organisations budget for AI. Procurement compares price lists. Cost is decided by run-time behaviour: which effort level is on, how long an agent keeps working, and whether a failed request is still billed.

In this test it was billed. The token allowance was spent and nothing came back. Think of a taxi with a lower per-mile rate that still charges you when the driver circles the block.

The person who feels this first is the developer or analyst waiting 20 minutes for an answer that never arrives. The invoice follows later.

Questions to put to your team

Ask which reasoning or effort level each production workflow uses, and who chose it. Ask whether any workload runs at the highest setting by default.

Ask whether spend is tracked per request, including requests that returned no usable result. A monthly total hides them.

Ask how much of your input volume is cached. Opus 5.5's cache-read price fell 60%, and Willison says 90%+ of input tokens in long agentic conversations are billed at cached prices.

Finally, ask whether cheaper models could handle simpler tasks. Willison reports moving a demo to GPT-6 Luna, and says it is fast and competent at SQL queries and building web pages.

Rates on the newest mid-tier models are dropping. Budget control now depends less on the rate card and more on how the meter is set.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Simon Willison.

Share this insight