Skip to main content

Cloudflare launches Clef-omni and cuts Clef-flash pricing

11 OCTOBER 2026·2 MIN READ·2 SOURCES

On October 9, 2026, Cloudflare made Clef-omni available on Workers AI. It joins Clef and Clef-flash in Cloudflare’s open-weight decision-model family and accepts text, images, audio and video. The company also cut Clef-flash’s input-token price and reduced its hosted context window.

Cloudflare launches Clef-omni and cuts Clef-flash pricing

Key takeaways · 4

  • 01

    Clef-omni is available on Workers AI and accepts text, images, audio and video.

  • 02

    Cloudflare says Clef-omni aligns video soundtracks with frames and handles both together.

  • 03

    Clef-omni is priced at $0.15 per million input tokens, with a 64,000-token context window.

  • 04

    Clef-flash now costs $0.038 per million input tokens, but its hosted context window is 24,000 tokens.

One request across four input types

Clef-omni joins Clef and Clef-flash in Cloudflare’s open-weight decision-model family.[1] It accepts text, images, audio and video, and Cloudflare says it processes each modality directly in one request.[1] The company says the model aligns a video’s soundtrack with its frames, allowing it to reason over what is heard and seen simultaneously.[1] Each request can include up to four audio clips and two videos.[2] Unlike a text-generating model, Clef-omni scores allowed answers.[1]

Price, context and reported speed

Cloudflare lists Clef-omni with a 64,000-token context window and a price of $0.15 per million input tokens.[1] The company reports about 20 milliseconds for text requests and under 100 milliseconds for image or audio inputs.[1] It reports about 300 milliseconds to process a 21-second video clip with sound.[1] Those figures are Cloudflare’s reported processing times; the evidence does not specify how they vary across workloads.[1]

Clef-flash gets a lower hosted price

Cloudflare cut Clef-flash’s input-token price from $0.090 to $0.038 per million and says the new price is below Jev’s.[1] At the same time, it reduced the hosted context window from 64,000 to 24,000 tokens.[1] Cloudflare says 0.24% of Clef-flash requests exceed the new 24,000-token limit.[1] Its published weights are unchanged, and Cloudflare says self-hosting can support a 256,000-token context window.[1] Teams comparing cost should therefore also check whether their workloads fit the hosted limit.[1]

Serving changes for Clef

Cloudflare also says it optimized Clef’s serving on Workers AI so the model returns decisions up to twice as fast, without changing its weights.[1] The company lists scores of 94.8 on BANKING77 (macro-F1) and 97.7 on CLINC150+OOS (macro-F1).[1] These figures describe the listed benchmark rows; they do not, on their own, establish how a model will perform on a particular team’s data or workflow.[1]

Teams evaluating decision models can compare Clef-omni’s multimodal inputs and listed price with the constraints of their own workloads. Clef-flash’s price cut comes with a shorter hosted context window, so cost comparisons should account for input length as well as token rates.

Why it matters
Story quiz

Test yourself on this story — 1 question.

Create a free account to take the quiz, earn XP, and get a daily session built for your industry.

Take the quiz

How this developed

  1. 11 October 2026

    Cloudflare launches Clef-omni and cuts Clef-flash pricing

Sources

AI fluency, one session a day, built for your work.