- Cloudflare released Clef and Clef-flash, open decision models that return typed answers with probabilities. Benchmark claims in the announcement are Cloudflare's own.
- Cloudflare says Clef classified a domain in 2.2 seconds, against 4.7 for an LLM. The key governance question is who sets the point where software defers to a person.
- Before adopting, test on your own past cases, assign an owner for the confidence cutoff, and check how fine-tuning data is handled.
The decision moves out of the chat window
Many jobs that companies give to AI are not writing jobs. They are sorting jobs. Is this message urgent? Which team owns it? Is this website safe?
On October 1, Cloudflare released two models built for that kind of work: Clef and Clef-flash. Cloudflare calls them decision models. The idea is simple. Instead of writing a paragraph, the model picks from a fixed set of answers and says how likely each one is.
This shows a shift in where human judgment sits. When a model returns odds, the important decision is no longer the answer. It is the cutoff that tells the software when to act and when to ask a person.
What Cloudflare released
Both models run on Cloudflare's Workers AI platform. Teams that prefer to keep things in-house can instead download the weights from Hugging Face. The license there is Apache 2.0, a permissive open-source license.
The models follow the API of Jev, a decision model from Typesafe AI. Cloudflare says that makes switching easy.
Cloudflare says Clef ranks first on the Jev Decision Index. It also says Clef beat Jev in three of four areas on Typesafe's own evaluation suite. These are Cloudflare's own test runs. The source offers no independent check.
How a decision model works
A standard large language model (LLM) writes its answer one word at a time. That is flexible, but slow, and the result is free text that code must then interpret.
Cloudflare describes a different path. Clef reads the input once. It then scores every allowed answer in parallel. Because it generates no text along the way, Cloudflare says it runs faster than models that write step by step.
Cloudflare built Clef on open Qwen models. It kept the base models frozen and trained small add-on layers on top. One training goal uses a Brier loss, a measure that penalizes probabilities that do not match real outcomes. In plain terms, Cloudflare is trying to make a stated 90% mean something.
Cloudflare also says Clef accepts images, which it says Jev does not. Its context window is 64k, against 32k for Jev. The context window is how much input the model can read at once.
A test from Cloudflare's own security team
Cloudflare's Threat Intelligence team has been testing Clef to classify website domains. Cloudflare gives an example in which Clef might score a domain at 95% fashion site, 85% ecommerce and under 1% phishing.
Fetching, rendering and classifying a site took Clef 2.2 seconds. The same workflow took 4.7 seconds with gpt-oss-120b, and that LLM returned only two classifications.
This is one internal workflow, tested by the vendor. It is a useful example, not a general measure of speed.
Who owns the cutoff
Cloudflare presents the output as an input to other code. Software can use the odds to send a support ticket to a team, raise an escalation, or pass the case to a person.
The company's argument is that many agent decisions can now run without a person reviewing each one. That puts weight on the rule that decides which cases still reach a person.
Think of a credit score with an approval line. The score is only a number. The policy that says "approve above this, review below this" is where the judgment sits. A decision model works the same way.
A team that sets the line too low lets more wrong calls pass automatically. A team that sets it too high sends work back to people and gives up the savings. Either way, someone in the organization must own that number.
Fine-tuning brings a data question
Cloudflare also announced a fine-tuning service for Clef. It starts with Cloudflare's forward-deployed engineers working alongside customers. A self-serve platform is planned for later. Cloudflare says it is still early.
The source notes a tradeoff. A tuned model may give up general performance for accuracy in one domain. The pipeline also uses AI Gateway to capture a customer's AI traffic as a dataset. Cloudflare says it does not read, store or train on requests unless a customer chooses its fine-tuning product. Leaders should check how that choice is handled in their contract.
Questions to put to your team
First, list the AI calls in your workflows that really just pick a label. Those are candidates for a decision model.
Second, ask who sets the confidence line for handing work to a person, and who reviews it.
Third, test on your own past cases. A vendor index does not show how well the stated odds match your outcomes.
Fourth, ask where logged requests go if you tune a model on your data.
A model that answers with odds gives the real decision back to whoever picks the cutoff.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Cloudflare.





