Reported savings, internal use: Up to 30% (Source: Cloudflare blog (September 30, 2026))
Model choice is becoming a company decision
Today, most employees using AI coding tools pick their own model. Cloudflare says that often means they pick one that is overkill. Its example: you do not need Opus-level intelligence to summarize an email.
Cloudflare's answer is to take the choice away from the user. This is the idea worth noticing. Once a gateway decides which model answers each request, model choice stops being a personal habit. It becomes a policy that someone has to own, inspect and defend.
What Cloudflare released
On September 30, Cloudflare put its Auto Router into public beta. It runs inside AI Gateway, the service that every AI request from a company's users, agents and tools can pass through. Staff set their model to cloudflare/auto. The router then sends each request to a model it judges capable enough for the task.
Cloudflare reports up to 30% savings in its own OpenCode deployment, compared with using only frontier models such as OpenAI Sol and Anthropic Claude Opus. The router is free during the beta.
How the router decides
The design has two stages. First, the router builds a pool of eligible models. Models that cannot handle the type of request are dropped. So are providers that are down. The pool also reflects the credentials, access policies and spend limits attached to the gateway.
Second, a classifier reads the most recent messages. It runs on Workers AI, Cloudflare's GPU service at the edge of its network. The classifier assigns probabilities across 14 task categories, such as coding, planning and research. It also grades each request on a five-point scale for four traits: how complex it is, how vague it is, how much rides on it, and how much it leans on the earlier conversation.
A separate table then pairs those readings with benchmark results for each model. The output is an estimate of which model suits the job. Price comes last. On simple requests, cost counts for a lot, so a smaller model can win. The harder the request looks, the less price matters, and the more likely a stronger model is picked.
One benefit is legibility. Cloudflare says you can inspect each request's predicted category and complexity to see why a model was chosen. Adding a new model also needs no retraining, only new benchmark-derived weights.
The hidden cost: switching models
Long agentic sessions, such as debugging, build up a large context. Models cache that context so they do not reread it at full price. Switching models throws the cache away. The new model must write the whole context again.
Cloudflare says the router prices this in. Within a turn, it tends to stay on one model. Across turns, it adds a switching penalty that grows with the context already built up. Cloudflare also notes that most models cannot read another model's reasoning tokens, so a switch may force the new model to redo that work at output prices.
The practical lesson is that a cheaper list price does not guarantee a cheaper outcome. Cloudflare says a model that looks cheaper may use far more tokens to finish the job.
What the numbers show, and what they do not
Cloudflare tested the router on its own general knowledge work benchmark. It uses simulated tools for email, calendars, Slack, files, travel and finance. Cloudflare ran each model three times on each of 97 tasks. It compared the router with OpenAI's GPT-6 Sol and Anthropic's Claude Opus 5.5.
It reports similar performance at 80% of Sol's cost and 35% of Opus's cost.
Treat these as a vendor's own results. The benchmark, tasks and tools are Cloudflare's. The 30% figure comes from a different measure: internal OpenCode usage. The two should not be added together or read as one result. Cloudflare says the router works best across a wide mix of knowledge work, and that savings grow with the volume of non-frontier work.
The questions this raises
The classifier does rate how high the stakes are for each request. Cloudflare does not say how that rating changes the choice of model. Its only stated rule is that price matters less as difficulty rises.
That matters because of Cloudflare's own example. It says you would not want to block a top model from your security engineering team. The source does not show that the router guarantees this access. Leaders should treat it as an open question, not an assurance.
A gap also remains for regulated data. Cloudflare lists zero-data-retention filtering as a near-term goal, not a current feature. It also lists provider capacity and reasoning-level selection as future work. Until then, the pool rules rest on what you have configured at the gateway.
What leaders should ask
Ask who owns the routing policy. If a gateway chooses models, someone must be accountable for what it chooses.
Ask Cloudflare how the stakes rating changes model choice. Ask for evidence that high-stakes work, such as security engineering, reaches a top model.
Ask for your own test. Run a sample of your real tasks through the router and a frontier model, and compare quality and total cost, not list prices.
Ask which data may go to which provider. Confirm that gateway access policies cover your retention requirements before sensitive work is routed.
Ask whether you can see each decision. Cloudflare says the category and complexity behind each choice can be inspected. Make that part of any review.
The line to remember: when software picks the model, the choice does not disappear. It moves to whoever sets the rules.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Cloudflare.





