Skip to content
Innovation & Growth

Cloudflare's cheaper AI decision model can read less text per request

Clef-flash now costs $0.038 per million input tokens. The hosted version's input limit fell from 64k to 24k tokens.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Cloudflare's cheaper AI decision model can read less text per request
In brief
  • Cloudflare cut Clef-flash to $0.038 per million input tokens and launched Clef-omni, which accepts audio and video. It also made hosted Clef faster.
  • The cheaper hosted Clef-flash now accepts 24k tokens per request, down from 64k. Cloudflare says only 0.24% of its requests exceeded 24k.
  • Before switching, test your own input sizes, check the speed and accuracy claims on your data, and decide who owns the decision thresholds.

Every price cut has a design choice behind it. When a vendor makes automated decisions cheaper, the useful question is what the product now does less of. Cloudflare's latest update to its Clef models shows the pattern clearly, because the company says openly what it traded away.

What Cloudflare announced

On October 9, Cloudflare released Clef-omni. The new model takes audio, video, images and text in one call. Cloudflare also made hosted Clef faster and cut the price of Clef-flash. Clef-omni costs $0.15 per million input tokens. Clef stays at $0.24.

$0.038
Clef-flash price per million input tokens
Source: Cloudflare, down from $0.09 (October 9, 2026)

Cloudflare says Clef-flash is now cheaper than Jev, a decision model from TypeSafe. The post gives no Jev price, so readers cannot check that comparison from the source. Cloudflare points customers to its developer docs for current pricing.

The trade behind the discount

To reach the lower price, Cloudflare cut the hosted Clef-flash context window from 64k tokens to 24k. The context window is how much input the model can read in one request. A token is a small chunk of text, roughly a word or part of one.

0.24%
Requests above 24k input tokens
Source: Cloudflare usage data (October 9, 2026)

Cloudflare says this figure justified the cut. The number comes from Cloudflare's own traffic. Your documents, call recordings or contracts may be longer than its typical request.

The cut applies to the hosted service only. Cloudflare says the open weights on Hugging Face are unchanged and were trained for a 256k window if you self-host. Cloudflare tells customers with larger inputs to move to Clef, which keeps 64k.

How a decision model differs from a chatbot

Clef models are not large language models. They do not write text back. They pick from options you define, such as spam or not spam, and attach a confidence score.

Cloudflare describes the Clef-omni mechanism this way. It starts from Qwen3-Omni-30B-A3B-Instruct, an existing open model, and keeps its comprehension layers. It discards the speech-output parts. It freezes that base and trains small add-on layers called low-rank adapters (LoRA). It then tunes the scores so that stated confidence matches how often the model is right, using a method called Brier score calibration.

At run time, the model reads the whole input once and scores every option at once. It skips transcribing audio or captioning video first. By Cloudflare's timing, half of text-only requests finish within roughly 130 milliseconds. A 21-second video clip with its sound track is scored in about one and a half seconds, in a single call.

130 ms
Median text decision time, Clef-omni
Source: Cloudflare (October 9, 2026)

Cloudflare measured these timings itself. The post also mentions benchmark results, but the text does not include the numbers.

Why the shift matters

Cloudflare describes teams using Clef today. A public docs repository closes spam issues. Its EmDash content system screens plugin libraries for phishing. A data loss prevention team scans for government IDs. A threat intelligence team looks for malicious domains.

Cloudflare argues that classification used to need a specialised machine learning team and a training corpus. Now it is an API call. That claim comes from the vendor, and it is plausible. It also moves a judgment from a team you can interview to a score you must interpret.

The lesson here is that cheap decisions spread fast, and risk spreads with them. A spam filter that errs costs little time. A data-loss or malicious-domain filter that errs has different costs. Cloudflare's post does not say which thresholds its own teams use.

Cloudflare also says Clef is Jev-API compatible and works by changing the model ID. That makes trying it easy. It also makes it easy for a team to adopt without review.

Questions to put to your team

First, how long are our real inputs? Measure them before assuming 24k is enough. Anything longer needs Clef at $0.24 or self-hosting.

Second, what happens to an input that exceeds the limit? Find out whether the request fails, is cut short or is routed elsewhere, and who sees it.

Third, have we tested it on our own labelled examples? Vendor benchmarks and speed figures do not replace a trial on your data.

Fourth, who sets the confidence threshold for action, and who reviews the cases below it? A calibrated score is only useful if someone owns the cut-off.

A lower price is welcome. It is worth reading the new limit with the same care as the new number.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Cloudflare.

Share this insight