Skip to content
Brief The AI-First Web ·

Cloudflare adds audio and video model Clef-omni, cuts Clef-flash price

Cloudflare also sped up hosted Clef and shrank the Clef-flash context window to pay for the lower price.

In brief
  • Cloudflare launched Clef-omni, which takes text, images, audio and video in one call, at $0.15 per million input tokens.
  • Clef-flash now costs $0.038 per million input tokens, but its hosted context window fell from 64k to 24k.

Cloudflare said in a blog post on 9 October that it launched Clef-omni, a decision model that takes text, images, audio and video. It costs $0.15 per million input tokens. Cloudflare also cut Clef-flash from $0.09 to $0.038 per million, which it says is now below Jev. Clef stays at $0.24. It said hosted Clef is faster after a move to SGLang, a tool for serving models.

Cloudflare paid for the Clef-flash cut by shrinking its hosted context window, the input it can read at once, from 64k tokens to 24k. It said only 0.24% of requests exceed 24k. The Hugging Face weights are unchanged and were trained for 256k. Clef-omni scores the allowed options in one pass and writes no text, so Cloudflare says it needs no transcription or captioning step. Cloudflare reports about 130 ms for text at the median and about 1.5 seconds for a 21-second video. Its post points to charts of new Clef speeds and benchmarks; those figures are not reproduced here. The post body does not state Jev's price.

The cheaper Clef-flash trades capacity for price, so buyers should read the limits beside the price. Useful questions: what share of our inputs exceed 24k tokens? Which pipelines could drop a transcription stage? Do Cloudflare's latency figures hold on our data?

A WebPulse Brief: a short report of an important event, written by the WebPulse Newsroom with AI assistance and checked against the reporting below. When there is more to explain, we follow up with a full story. How we use AI.

Reporting: Cloudflare.