Cloudflare launches Clef-omni and cuts Clef-flash pricing
On October 9, 2026, Cloudflare made Clef-omni available on Workers AI. It joins Clef and Clef-flash in Cloudflare’s open-weight decision-model family and accepts text, images, audio and video. The company also cut Clef-flash’s input-token price and reduced its hosted context window.

Key takeaways · 4
- 01
Clef-omni is available on Workers AI and accepts text, images, audio and video.
- 02
Cloudflare says Clef-omni aligns video soundtracks with frames and handles both together.
- 03
Clef-omni is priced at $0.15 per million input tokens, with a 64,000-token context window.
- 04
Clef-flash now costs $0.038 per million input tokens, but its hosted context window is 24,000 tokens.
One request across four input types
Clef-omni joins Clef and Clef-flash in Cloudflare’s open-weight decision-model family.[1] It accepts text, images, audio and video, and Cloudflare says it processes each modality directly in one request.[1] The company says the model aligns a video’s soundtrack with its frames, allowing it to reason over what is heard and seen simultaneously.[1] Each request can include up to four audio clips and two videos.[2] Unlike a text-generating model, Clef-omni scores allowed answers.[1]
Price, context and reported speed
Cloudflare lists Clef-omni with a 64,000-token context window and a price of $0.15 per million input tokens.[1] The company reports about 20 milliseconds for text requests and under 100 milliseconds for image or audio inputs.[1] It reports about 300 milliseconds to process a 21-second video clip with sound.[1] Those figures are Cloudflare’s reported processing times; the evidence does not specify how they vary across workloads.[1]
Clef-flash gets a lower hosted price
Cloudflare cut Clef-flash’s input-token price from $0.090 to $0.038 per million and says the new price is below Jev’s.[1] At the same time, it reduced the hosted context window from 64,000 to 24,000 tokens.[1] Cloudflare says 0.24% of Clef-flash requests exceed the new 24,000-token limit.[1] Its published weights are unchanged, and Cloudflare says self-hosting can support a 256,000-token context window.[1] Teams comparing cost should therefore also check whether their workloads fit the hosted limit.[1]
Serving changes for Clef
Cloudflare also says it optimized Clef’s serving on Workers AI so the model returns decisions up to twice as fast, without changing its weights.[1] The company lists scores of 94.8 on BANKING77 (macro-F1) and 97.7 on CLINC150+OOS (macro-F1).[1] These figures describe the listed benchmark rows; they do not, on their own, establish how a model will perform on a particular team’s data or workflow.[1]
Teams evaluating decision models can compare Clef-omni’s multimodal inputs and listed price with the constraints of their own workloads. Clef-flash’s price cut comes with a shorter hosted context window, so cost comparisons should account for input length as well as token rates.
Why it matters
Test yourself on this story — 1 question.
Create a free account to take the quiz, earn XP, and get a daily session built for your industry.
Take the quizHow this developed
11 October 2026
Cloudflare launches Clef-omni and cuts Clef-flash pricing
Sources
- Clef-omni adds audio and video input, Clef-flash is now cheaper, and Clef is faster · Changelogdevelopers.cloudflare.com
- clef-omni (Cloudflare) · Cloudflare AI docs · Cloudflare AI docsdevelopers.cloudflare.com