Skip to content
The AI-First Web

Anthropic's Haiku 5.5 cuts short-prompt prices 90%, making agents cheaper

WebPulse's view: cheaper Haiku 5.5 weakens cost as a brake on agent numbers. Oversight must fill the gap.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Anthropic's Haiku 5.5 cuts short-prompt prices 90%, making agents cheaper
In brief
  • Anthropic priced Claude Haiku 5.5 at $0.10 input and $0.50 output per million tokens for prompts up to 100K tokens, against $1 and $5 for Haiku 4.5.
  • WebPulse's view: cheaper small models make cost a weaker brake on how many agents a company runs, so oversight has to fill the gap.
  • Leaders should set permissions, logging and spending rules for small-model agents, and test the price tiers against their real prompt sizes.

WebPulse's view: the price of a model call has long worked as a quiet brake on automation. Teams asked whether a task was worth the money before they handed it to software. When the per-token price of short prompts falls tenfold, and Anthropic estimates the average saving at about 75%, that question filters far less. A different one gains weight: what are we willing to let software do, unattended, at volume?

What Anthropic changed

Anthropic released Claude Haiku 5.5 on October 7, 2026. The company places it at the small, quick end of its Claude 5.5 lineup. Anthropic says it is built for high-volume, cost-sensitive work such as summarization, classification and request routing.

The price depends on how long the prompt is. A token is a small chunk of text, roughly a word fragment. Short prompts, up to 100K tokens, cost $0.10 per million tokens going in and $0.50 coming out. Longer prompts cost $0.50 and $2.50. Haiku 4.5 cost $1 and $5.

$0.10 / $0.50
Haiku 5.5 price, prompts up to 100K tokens (input / output, per million tokens)
Source: Anthropic, Claude Haiku product page (October 7, 2026)

Why the headline number needs care

The New Stack reported that Haiku 4.5 had one flat rate for any prompt length. Haiku 5.5 does not. The tenfold cut applies to short prompts only. Long prompts get a cut of 50%.

Most traffic sat on the short side before. The New Stack says Anthropic reports roughly nine in ten Haiku 4.5 requests would land in the cheaper band. Anthropic estimates the blended saving at about 75%.

That average is lower than 90% for two reasons. About one in ten requests are long prompts, and those get only the smaller 50% cut. Haiku 5.5 also uses a new tokenizer, the part that cuts text into tokens. It produces slightly more tokens per task, so the same work is billed as more units.

About 75%
Average savings versus Haiku 4.5, as estimated by Anthropic
Source: Anthropic, as reported by The New Stack (October 7, 2026)

Other levers exist. Anthropic lists discounts of up to 90% for prompt caching, which reuses repeated input, and 50% for batch processing. Haiku 5.5 also adds a setting no earlier Haiku offered: effort controls, a dial for how much work the model puts into each task. Teams can use it to trade cost against intelligence. The New Stack says the default is medium.

Why this is more than a discount

Anthropic does not present Haiku 5.5 as a text tool alone. It pitches the model for repetitive screen work: filling in forms, keying in data and shifting information from one application to another. It also suggests a pairing. A smarter model plans the job and passes pieces to Haiku, so many agents can work side by side.

The capability jump is large on Anthropic's own numbers. On a test of operating a computer, the offline part of OSWorld 2.1, Anthropic's evaluations give Haiku 5.5 a score of 72.4%, per The New Stack. Haiku 4.5 managed 15.7%. OpenAI's GPT-6 Luna got 48.9%. These are vendor-reported figures, and Sonnet 5.5 still ranks higher.

72.4% vs 15.7%
OSWorld 2.1 (offline subset) score, Haiku 5.5 vs Haiku 4.5
Source: Anthropic evaluations, as reported by The New Stack (October 7, 2026)

Here is the argument. A capable agent that costs a fraction of before is easier to approve and easier to multiply. The limit then moves from money to permissions, logging and review. Those controls take real effort, and no price sheet supplies them.

This is WebPulse's interpretation, not a finding in the sources. Anthropic says it tested Haiku 5.5 against its safety, security and reliability standards and discusses results in a system card. The sources do not show how agents behave inside a customer's own systems. That part is for the buyer to govern.

What the sources leave open

Haiku 5.5 is not the cheapest choice. The New Stack points to Alibaba's Qwen3.7 Flash at $0.03 and $0.13 per million tokens on its international service. That rate covers inputs up to 32,000 tokens.

On one knowledge-work benchmark, GDPval-AA v2.1, Artificial Analysis gives Z.ai's GLM-5.3-Flash 1,647. Anthropic reports 1,620 for Haiku 5.5. Anthropic compares the model only with its own and GPT-6 Luna.

Price and scores also leave out the cost of mistakes. A cheaper model that errs more can cost more once people repair its output.

Questions to put to your team

First, which of our workloads stay under 100K tokens, and which cross it? Input price rises fivefold at that line, so prompt length now shapes the bill.

Second, what does the tokenizer change do to our real costs? Anthropic's 75% is a blended estimate, not a promise for any one workload.

Third, if an agent can fill forms and move data between apps, what can it reach? List the accounts, data and actions it may take, and name who reviews its logs.

Fourth, who sets the effort level for each task, and who approves a change? A setting that moves both cost and quality should not rest with one developer.

Cheaper intelligence does not remove the need for judgment. It moves that judgment from the purchase order to the permissions list.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Anthropic.

Share this insight