Cloudflare said in a blog post on 9 October that it launched Clef-omni, a decision model that takes text, images, audio and video. It costs $0.15 per million input tokens. Cloudflare also cut Clef-flash from $0.09 to $0.038 per million, which it says is now below Jev. Clef stays at $0.24. It said hosted Clef is faster after a move to SGLang, a tool for serving models.
Cloudflare paid for the Clef-flash cut by shrinking its hosted context window, the input it can read at once, from 64k tokens to 24k. It said only 0.24% of requests exceed 24k. The Hugging Face weights are unchanged and were trained for 256k. Clef-omni scores the allowed options in one pass and writes no text, so Cloudflare says it needs no transcription or captioning step. Cloudflare reports about 130 ms for text at the median and about 1.5 seconds for a 21-second video. Its post points to charts of new Clef speeds and benchmarks; those figures are not reproduced here. The post body does not state Jev's price.
The cheaper Clef-flash trades capacity for price, so buyers should read the limits beside the price. Useful questions: what share of our inputs exceed 24k tokens? Which pipelines could drop a transcription stage? Do Cloudflare's latency figures hold on our data?