Skip to content
The AI-First Web

AI bills are set by how apps are built, InfoWorld's five fixes argue

Model choice, caching, context size and output limits decide what each AI request costs, per Matthew Tyson

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
AI bills are set by how apps are built, InfoWorld's five fixes argue
In brief
  • InfoWorld's Matthew Tyson argues AI spend is shaped by five engineering choices: model routing, two kinds of caching, trimmed context and capped output.
  • The deeper problem he names is attribution: leaders often cannot say where AI money goes or what value it returns.
  • Ask your team which of the five levers are in use and whether spend is tracked per product or tenant.

The AI bill is decided before the invoice arrives

Finance teams spent a decade learning to read the cloud bill. InfoWorld contributor Matthew Tyson writes that generative AI adds a faster, more opaque layer on top of it. His sharpest point is not the size of the bill. It is attribution: where exactly the money goes, and what value it delivers to the business.

The lesson here is that an AI bill is mostly an engineering outcome. Every request carries choices about which model answers, what context travels with it and how long the reply runs. Those choices are made in code, long before anyone in finance sees a number. Leaders who ask only about vendor price are asking about the smaller part.

Lever one: match the model to the task

Tyson says teams start prototypes on the most powerful model so they can avoid its limits. Production is a different job. He argues the goal is the least powerful model that does the work.

He describes a routing layer that sends simple jobs, such as classification, basic text parsing and intent detection, to cheaper models. He names RouteLLM and Semantic Router for this, and GPT-4o mini and Claude 3 Haiku as examples of the cheaper models.

AI gateways such as Kong, Cloudflare AI Gateway and Portkey go further. Tyson says they pull usage data into one place. That lets a team enforce a hard token limit for each customer or each service. He also notes that gateways add complexity and cost of their own.

Levers two and three: stop paying twice

Ordinary caching matches exact text. People phrase the same question many ways, so exact matching rarely hits. Semantic caching instead converts each prompt into numbers that capture its meaning, then looks for a close match among past answers. Tyson gives a similarity threshold of 0.92 as an example. A hit skips the model entirely.

He lists the trade-offs. The lookup itself costs something. A threshold set too low can serve recycled answers to nuanced questions. He calls the method practically useless for open-ended creative work, but strong for support bots and internal knowledge bases.

Prompt caching works differently. The provider keeps the large background material you send with each question, so you stop resending it. Tyson says cached tokens get a discount of 50% to 90%.

50% to 90%
Discount on cached prompt tokens
Source: InfoWorld, Matthew Tyson (October 7, 2026)

The details differ by provider. Tyson says OpenAI's caching is automatic, with a 1,024-token minimum and an exact match from the first token. A timestamp or user ID placed at the top of the instructions breaks the match, and the full price returns. He says Anthropic and Google use explicit caching that engineers must build in, sometimes with an upfront write premium.

Levers four and five: send less, ask for less

Today's models can read a very large prompt, in the range of one to two million tokens. Tyson sees the temptation to paste in whole codebases or years of logs as a finops blunder. He also warns that accuracy can fall as the payload grows, a problem he calls "context rot".

His fix is reranking. Cast a wide net of 50 candidate passages from a cheap search, then use a small scoring model to keep the top three or four. Tyson says this adds about 100 milliseconds and routinely cuts prompt tokens by 80% or more. These are his figures, not a measured benchmark.

80% or more
Prompt token reduction from reranking, per the author
Source: InfoWorld, Matthew Tyson (October 7, 2026)

On output, Tyson says generated tokens are priced 3x to 5x higher than input tokens. Every pleasantry is paid for at the higher output rate. He recommends a hard cap on length as a safety stop, stop sequences that end generation early, and structured output formats.

3x to 5x
Output token price versus input, per the author
Source: InfoWorld, Matthew Tyson (October 7, 2026)

He adds a caution. Forcing a model to answer in a single word, with no room to reason, can cut accuracy sharply on complex reasoning tasks. The aim is that every paid output token does real work.

Questions to put to your team

Start with attribution. Can the team show AI spend by product, tenant or service? If not, a gateway with per-tenant budgets is the first thing to scope, weighed against its own cost.

Then check the basics. Which requests go to the most powerful model, and do they need to? Are stable instructions placed first so provider caching can match them? Is a cache or reranker suited to the workload, such as a support bot, or would it flatten nuanced answers?

One caution on the source. This is an opinion and how-to piece. It offers no customer data on savings, so treat its percentages as the author's claims and test them on your own traffic.

The cheapest token is the one you never send. Ask who in your company is responsible for sending fewer.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: InfoWorld.

Share this insight