- Cloudflare released Clef, a model that returns classifications with probabilities, so agent code can act alone or hand a case to a person.
- The human role shifts from reviewing each case to setting the confidence level at which software may act alone.
- Benchmarks are Cloudflare's own. Ask who sets thresholds, how calibration is tested, and what data fine-tuning requires.
The human does not leave. They move to the threshold.
Picture a customer support message arriving. In many workflows, someone reads it, judges how urgent it is and chooses a team. Cloudflare says its new Clef model can take that step in code.
In its launch post, Cloudflare says a human "does not necessarily need to be in the loop" for agentic decisions anymore. The model returns typed answers with probabilities. Code then routes the ticket, escalates it, or defers to a person.
This shows where the human job is going. People stop reviewing each case. They set the probability at which software may act alone. That is a policy decision about risk, and it sits in a setting that can easily go unreviewed by executives.
What Cloudflare released
Cloudflare released two decision models, Clef and the smaller Clef-flash. They run on its Workers AI service. The weights are also open under an Apache 2.0 license on Hugging Face. The API is compatible with Jev, a decision model from TypeSafe AI, so customers can swap one for the other.
Cloudflare lists two differences from Jev. Clef accepts images, while Jev handles text only. Clef also has a 64,000-token context window, double Jev's 32,000. A token is a small chunk of text the model reads.
How it works
A chat-style AI model writes its answer one word at a time. That is slow, and the wording can vary. Clef does not write text.
Both models start from an existing Qwen model: Qwen3.8-27B underpins Clef, and Qwen3.5-9B underpins Clef-flash. Cloudflare leaves that base untouched. One pass through it digests the input, and then every permitted answer is scored at the same time. Because no text is written along the way, Cloudflare says Clef runs much faster than chat-style models.
What Cloudflare did train were small add-on components. It paired that training with a Brier loss, a measure meant to make stated probabilities match real outcomes. It also uses its own variant of RLCD (Reinforcement Learning for Calibrated Decisions), a method TypeSafe used to train Jev, according to The Decoder. That shared method name and the API compatibility together support Cloudflare's pitch that one model can replace the other.
Calibration is the key idea. A 90% answer should be right about nine times in ten. Without that, the number is decoration.
What the evidence does and does not show
The benchmarks are Cloudflare's own. Cloudflare says Clef leads the Jev Decision Index and beat Jev in three of four areas on TypeSafe's evaluation suite. Cloudflare also says its models were faster than the other decision models across 43 benchmarks, except Laya, which it says trades quality for speed.
The 2.2-second versus 4.7-second result is one workflow: Cloudflare's threat intelligence team classifying domains. It is a useful example. It is not a general speed comparison.
The material reviewed gives no calibration results from independent testers. That is the number a risk owner needs most.
Who carries the risk
Cloudflare names where it wants to use fine-tuned versions internally: Trust & Safety submissions, support triage, and deciding whether a crawler is a good bot or a bad bot. Cloudflare describes these as planned uses, not reported results.
Each one affects someone outside the company. A customer waits, an account is flagged, a crawler is blocked. If the threshold is set too loosely or too tightly, those people are the first to be affected.
Fine-tuning also has a trade. Cloudflare says a tuned model may give up some general performance for accuracy in one domain. It offers a hands-on engineering service first, with self-serve tuning later. Cloudflare says it does not read, store or train on requests unless a customer uses fine-tuning.
Questions to put to your team
Which decisions in our workflows could a probability score replace? Who owns the confidence level for each one, and who approves a change?
How would we test whether a stated 90% is right about 90% of the time on our own data? What share of cases still goes to a person, and is that person trained for the hard ones?
If we fine-tune, which of our request data leaves the building, and under what terms?
A faster model does not remove judgment from the process. It compresses that judgment into one number, set once and applied thousands of times.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Cloudflare.





