Cloudflare Clef vs Clef-flash vs Clef-omni: which model gets each decision
Three Clef model IDs now sit on Workers AI, and each one wins a different kind of request. Three plain tests, asked in order, pick the right one before you write a line of code.
Cloudflare Clef vs Clef-flash vs Clef-omni: which one gets the call?#
Per Cloudflare's model pages on 10 October 2026, audio or video goes to Clef-omni, a state under 24,576 tokens to Clef-flash, and the rest to Clef. So deciding which Cloudflare Clef model to use takes three tests, and each is a fact you can check before you write code. First, look at the input and ask whether it carries audio or a video file with sound. Then count the tokens in the state, which is the text, JSON or record the model judges. Finally, multiply that count by the price.
A decision model is a different tool from a chatbot. You send a state and a list of typed questions. Then it returns a probability for every allowed option of every question, as the Clef model page puts it. Why does that shape matter? Our post on Jev versus large language models covers it.
So price is the third test, not the first. After all, a cheap model that cannot read the input is not a cheap answer. Cloudflare's 9 October 2026 launch post lists Clef at $0.24, Clef-omni at $0.15 and Clef-flash at $0.038 per million input tokens, down from $0.09.
Show data table
| Item | Value |
|---|---|
| Clef | 0.24 USD per 1M tokens |
| Clef-omni | 0.15 USD per 1M tokens |
| Clef-flash before 9 October | 0.09 USD per 1M tokens |
| Clef-flash now | 0.038 USD per 1M tokens |
Clef-flash now costs less than a sixth of Clef per input token.
Since all three models share one request shape, a wrong first pick costs you one string, not a rewrite. In practice, you can start on Clef-flash and move only the requests that fail a test.
What does a million decisions cost on each Clef model?#
At 3,400 tokens a decision, a million decisions cost $816 on Clef, $510 on Clef-omni and $129.20 on Clef-flash, at Cloudflare's prices read 10 October 2026. The bill scales with tokens per decision, so Clef-flash costs under a sixth of Clef at every state size it can hold. Also, output is not on the bill, since each model page lists only a price per million input tokens.
The sum is simple. Multiply tokens per decision by a million decisions, then by the price per million tokens. For example, say a support desk sends a state of about 3,400 tokens a million times a month. That is 3,400 million tokens, and at Cloudflare's 10 October 2026 price of $0.24, the bill is $816 on Clef. The same sum gives $510 on Clef-omni and $129.20 on Clef-flash.
For a reference point, Jev is TypeSafe AI's text-only decision model, and the Clef API is built to match it. TypeSafe AI's models page, read 10 October 2026, lists the Jev price per Mtok as $0.042. So at 3,400 tokens, Jev costs $142.80 per million decisions, and Clef-flash now undercuts it.
But the gap grows with the state. Under Cloudflare's same 10 October 2026 prices, a 16,000-token state costs $3,840 per million decisions on Clef and $608 on Clef-flash. Meanwhile, a short 800-token state costs $192 on Clef and $30.40 on Clef-flash.
Show data table
| Dimension | Clef | Clef-omni | Clef-flash | Jev |
|---|---|---|---|---|
| 800 tokens | 192 US dollars per million decisions | 120 US dollars per million decisions | 30.4 US dollars per million decisions | 33.6 US dollars per million decisions |
| 3,400 tokens | 816 US dollars per million decisions | 510 US dollars per million decisions | 129.2 US dollars per million decisions | 142.8 US dollars per million decisions |
| 16,000 tokens | 3,840 US dollars per million decisions | 2,400 US dollars per million decisions | 608 US dollars per million decisions | 672 US dollars per million decisions |
The bill scales with tokens per decision, so Clef-flash costs under a sixth of Clef at every state size it can hold.
Your monthly bill on each model
Put in your state size and monthly volume; the four prices are the ones Cloudflare and TypeSafe AI printed on 10 October 2026, and you can change them.
Monthly bill on Clef-flash
$129.20
- Monthly bill on Jev
- $142.80
- Monthly bill on Clef-omni
- $510.00
- Monthly bill on Clef
- $816.00
- Tokens left in the hosted Clef-flash window (0 means use Clef)
- 21,176
Arithmetic from Cloudflare and TypeSafe AI prices read 10 October 2026. Modelled, not measured.
So the money question is mostly a size question. If your state fits Clef-flash, the bill falls to less than a sixth of the Clef bill.
What changed for Clef, Clef-flash and Clef-omni on 9 October 2026?#
Cloudflare's 9 October 2026 post cut Clef-flash from $0.09 to $0.038 per million input tokens, cut its hosted window to 24k, sped up Clef and added Clef-omni. Three changes landed at once, and each one moves a different test in the routing rule. As a result, any page written about the 1 October launch now holds a stale price and a stale window.
First, the price cut came with a trade. Cloudflare's 9 October 2026 launch post states the cost plainly. Hosted Clef-flash "now has a context window of 24k, instead of 64k as previously advertised". But the weights on Hugging Face did not change. Per the same 9 October 2026 Cloudflare post, they were trained for a 256k window if you host them yourself.
Second, Clef got faster with no new weights. Cloudflare's 9 October 2026 post says most of the work was in the serving layer. That included a move to SGLang, an open source model server. In its 9 October 2026 table, Cloudflare puts the median for a state of about 800 tokens as falling from 262 ms to 152 ms.
262 ms
Before 9 October 2026
152 ms
Now
Cloudflare puts the median for a state of about 800 tokens as falling from 262 ms to 152 ms.
| Option | median latency in milliseconds |
|---|---|
| Before 9 October 2026 | 262 ms |
| Now | 152 ms |
Third, Clef-omni arrived as a new model, not an upgrade to Clef. Per Cloudflare's 9 October 2026 post, it is built on Qwen3-Omni-30B-A3B, a mixture-of-experts model. That means only some of its 30 billion parameters run on each token, about 3 billion. Cloudflare froze that base and trained low-rank adapters (LoRA), which are small add-on weights.
Which Clef model accepts audio, video and images?#
Clef and Clef-flash read text, JSON, images and video frames, while only Clef-omni takes wav or mp3 audio and mp4 or webm video, per Cloudflare's docs read 10 October 2026. Only audio, or a video file with its soundtrack, forces Clef-omni. Images and frame arrays do not.
This is the test most people get wrong, because "multimodal" sounds like one feature. In fact, the launch post says "With Clef, we supported images and video frame arrays". Also, the Clef-flash model page lists the same images field as Clef. So a photo of a damaged parcel can go to Clef-flash at the lowest price. The same holds for frames pulled from a camera feed.
But audio is the hard line. The Clef-omni model page, read 10 October 2026, adds an audio field of up to 4 clips and a videos field of up to 2. It also sets a condition on sound. Per that page, "A video's soundtrack is heard with its frames when every video in the request has one". As a result, a call recording, a voicemail or a clip where the sound matters has only one hosted home.
By contrast, Jev takes no media at all. Its models page is blunt: "No image, audio, or video input". So a Jev workload that already turns media into text upstream can keep doing that on any Clef model.
| Input | Clef | Clef-flash | Clef-omni | Jev |
|---|---|---|---|---|
| Text and JSON state | Yes | Yes | Yes | Yes |
| Images | Yes | Yes | Yes | No |
| Audio, wav or mp3, up to 4 clips | No | No | Yes | No |
| Video files, mp4 or webm, up to 2 | No | No | Yes | No |
How many tokens can a state hold on each hosted Clef model?#
Cloudflare's model pages, read 10 October 2026, list 65,536 tokens for Clef, 64,000 for Clef-omni and 24,576 for Clef-flash, whose open weights support 256k when self-hosted. A state over 24,576 tokens rules out hosted Clef-flash, and the docs say long text state is truncated rather than refused.
And that last part is the trap. Both the Clef and Clef-omni pages say it. "Long text state is truncated to fit the model's token limit". In other words, a state that is too long does not fail loudly. Instead, the model judges only part of it, and the answer may still look fine.
Show data table
| Item | Value |
|---|---|
| Clef | 65,536 tokens |
| Clef-omni | 64,000 tokens |
| Clef-flash | 24,576 tokens |
A state over 24,576 tokens rules out hosted Clef-flash.
By Cloudflare's own 9 October 2026 numbers, the Clef-flash 24k context window is enough for most traffic. Cloudflare's 9 October 2026 launch post says "only 0.24% of requests exceed 24k input tokens", and that was its reason for the cut. However, that figure is Cloudflare's own usage, not yours. So measure your own longest states before you switch.
The Clef-omni window needs one caveat. Its 64,000 figure comes from Cloudflare's model page, read 10 October 2026. The Hugging Face card gives the same default as a max length. Yet neither source says whether audio and video tokens count toward that window or sit beside it.
If a state runs past the 24,576 tokens on Cloudflare's 10 October 2026 Clef-flash page, you have two choices. Send it to hosted Clef, or host the Clef-flash weights yourself for the 256k window the launch post describes.
Where does Clef-omni score below Clef?#
In Cloudflare's own 9 October 2026 table, Clef-omni scores 60.2 on invoice exact actions against 64.7 for Clef, and trails Clef on every workflow row listed. Clef-omni buys modalities with accuracy, so a text decision sent to it costs less than Clef and scores lower.
These are vendor-run numbers on TypeSafe's workflow evals, printed in Cloudflare's 9 October 2026 launch post. Still, they are worth reading. They come from the people who built the model, and they still show a loss. On invoice processing, that 9 October 2026 Cloudflare table puts Clef-omni below Jev too, at 60.2 against 61.8.
| Workflow and metric | Clef-omni | Clef | Clef-flash | Jev |
|---|---|---|---|---|
| Invoice processing, exact actions | 60.2 | 64.7 | 57.1 | 61.8 |
| Invoice processing, primary action | 82.0 | 86.2 | 73.3 | 83.1 |
| Customer service, exact actions | 71.6 | 76.3 | 77.0 | 76.0 |
| Security incidents, exact actions | 61.7 | 62.9 | 61.7 | 61.7 |
| Agent trace observability, primary action | 65.8 | 68.5 | 69.8 | 71.6 |
Also, the table shows that Clef-flash is not simply a weaker Clef. On customer service exact actions, Cloudflare's 9 October 2026 table gives Clef-flash 77.0 and Clef 76.3. So the cheap model can win where the state is short and the labels are plain.
In short, send text-only work to Clef or Clef-flash. Pay the Clef-omni accuracy cost only when one call that hears and sees the input replaces a chain of models.
Why do the published scores swing so far between the three models?#
Cloudflare's 9 October 2026 table puts Clef-flash at 97.73 on home appliances but 66.77 on CLINC150+OOS, so a benchmark win says little about your own labels. Rankings flip by task, so the model is chosen on a sample of your own decisions, not on the table.
The same pattern runs through the rest of Cloudflare's 9 October 2026 benchmark table. For instance, Jev leads When2Call at 80.97 but trails at 52.27 on home appliances. Meanwhile, in that 9 October 2026 Cloudflare table, Clef-omni edges Clef on CLINC150+OOS, a test of sorting user requests into intents, at 97.7 against 97.43.
Show data table
| Clef-omni | Clef | Clef-flash | Jev | |
|---|---|---|---|---|
| Home appliances, case exact | 69.3 score | 82.95 score | 97.73 score | 52.27 score |
| When2Call, accuracy | 63.3 score | 72.37 score | 65.58 score | 80.97 score |
| CLINC150+OOS, macro-F1 | 97.7 score | 97.43 score | 66.77 score | 89.27 score |
| PhishNChips, accuracy | 73.2 score | 79.6 score | 75.05 score | 62.55 score |
| BANKING77, macro-F1 | 94.8 score | 94.2 score | 90.93 score | 79.74 score |
Rankings flip by task, so the model is chosen on a sample of your own decisions, not on the table.
So treat the table as a list of what to test, not as a ranking. Take a few hundred real decisions you have already labelled. Then run them through each model you are weighing, and compare the answers to your labels. Our post on what to settle before a System One model takes traffic covers how to set that test up.
What does an audio or video decision cost on Clef-omni?#
Cloudflare's Clef-omni page, read 10 October 2026, caps video at 256 input tokens a second, so a 21-second clip adds at most 5,376 tokens, while audio billing is unpublished. Video cost has a documented ceiling and audio cost has no published rate, so a media budget is an upper bound plus a measured trial.
In particular, the page says "Each second of video costs up to 256 input tokens". Cloudflare's page, read 10 October 2026, also caps a request at 2 videos of 60 seconds each. So on those limits, one full clip adds at most 15,360 tokens, and two clips add at most 30,720.
Show data table
| Item | Value |
|---|---|
| 21-second clip | 5,376 input tokens at most |
| one 60-second clip | 15,360 input tokens at most |
| two 60-second clips, the request maximum | 30,720 input tokens at most |
Video cost has a documented ceiling of 256 input tokens a second, so a budget is an upper bound.
But "up to" means the real count may be lower, and Cloudflare has not said how much lower. So the ceiling is a budget cap, not a forecast. In short, Cloudflare Clef-omni audio and video decisions carry one known cap and one unknown rate.
Audio is the open question. Cloudflare's 9 October 2026 launch post points to the docs for the answer. It promises "more detail on how image and audio modalities are converted to input tokens". Yet the Clef-omni page prints a token range for images and none for audio. The Workers AI pricing page, read 10 October 2026, prints none either. So run a small trial, read the usage object each response returns, and budget from what you measure.
On speed, Cloudflare's 9 October 2026 launch post is precise. It gives medians of about 130 ms for text and about 150 ms for an image. And the same 9 October 2026 Cloudflare post gives about 1.5 seconds for a 21-second video with sound.
Which Cloudflare docs do you build a Clef call from?#
Build from the Workers AI bindings page, then the Clef and Clef-omni model pages, then TypeSafe's System One API reference for the shared request shape. The official pages, read in build order, give the binding, the request shape, the media fields and the compatible API.
- Workers AI bindings: the AI binding in your Wrangler file, which exposes the models on
env.AI. - Clef model page: the
env.AI.runcall withmodel,stateandquestions, and the model IDs for Clef and Clef-flash. - Clef-omni model page: the
images,audioandvideosfields with their size and length limits. - TypeSafe System One API reference: the request shape Clef copies, with the three question types it supports.
What is the smallest Worker that sends each decision to the right Clef model?#
One TypeScript Worker can apply the rule: check for audio or video, estimate the state size, pick the model ID, then call env.AI.run once. Switching models is one model ID string, because all three share the same request shape.
First, the bindings page adds the binding to your Wrangler file as "ai": { "binding": "AI" }. Then the Worker below reads a request, picks a model and returns the answers. Its size check uses the character count of the state as a cautious stand-in for tokens. Once you have measured a sample, swap in your own token count.
export interface Env {
AI: Ai;
}
// Clef-flash hosted context window, per its model page
const FLASH_LIMIT = 24576;
export default {
async fetch(request, env): Promise<Response> {
const body = (await request.json()) as {
state: string;
audio?: string[];
videos?: string[];
};
// Step 1: audio or video forces Clef-omni.
const hasMedia = (body.audio?.length ?? 0) > 0 || (body.videos?.length ?? 0) > 0;
// Step 2: a long state rules out hosted Clef-flash.
const tooLong = body.state.length > FLASH_LIMIT;
const id = hasMedia ? "clef-omni" : tooLong ? "clef" : "clef-flash";
// Step 3: one call, the same shape for all three models.
const response = await env.AI.run(`@cf/cloudflare/${id}`, {
model: id,
state: body.state,
...(hasMedia ? { audio: body.audio, videos: body.videos } : {}),
questions: {
urgent: {
type: "noul",
instructions: "Is this support request urgent?",
},
team: {
type: "choice",
instructions: "Which team should handle this request?",
criteria: {
billing: "Payments, invoices, and refunds",
technical: "Outages, errors, and configuration",
sales: "Plans and upgrades",
},
},
},
});
return Response.json({ model: id, answers: response.answers });
},
} satisfies ExportedHandler<Env>; The question block comes from the Clef model page. So you can test the router before you write your own questions. Also, note the model field in the body. The Clef page accepts "clef" or "clef-flash" there, while the Clef-omni page accepts only "clef-omni". So the ID and the field must match.
When not to use Clef-omni, or any Clef model?#
Skip Clef-omni when the media can be turned into text once upstream, and skip every Clef model when the answer must be free-form text, because none of them generate any. A decision model returns probabilities over options you define, so a summary or a reply belongs to a text model instead.
Three cases point elsewhere. First, if you already transcribe calls for another reason, send the transcript to Clef and keep its higher scores. Second, if you need a written reply or a summary, use a large language model. After all, the Hugging Face card says "There is no free-form text generation and no output parsing". Our Jev versus LLMs post explains where each one fits.
Third, if your states run past the 65,536 tokens on Cloudflare's 10 October 2026 Clef page, no hosted Clef model reads them whole. In that case, host the Clef-flash weights for the 256k window. Or split the state into parts and decide on each one.
| Your case | Use instead |
|---|---|
| You already transcribe calls for another reason | Send the transcript to Clef and keep its higher scores |
| You need a written reply or a summary | A large language model; no Clef model generates free-form text |
| Your states run past 65,536 tokens | Host the Clef-flash weights for the 256k window, or split the state |
What should you read next on decision models?#
Earlier posts cover the basics this comparison skips: what a typed decision model is, the numbers to settle before launch, and the business case. The Cloudflare Clef vs Clef-flash vs Clef-omni comparison assumes the basics, which the sibling posts carry.
For the basics, read why a typed decision model beats a large language model on cost. Before launch, read the four numbers to settle for any System One model. For the business case, read why small AI calls on the page pay off. Then, once a model is chosen, running a multi-step AI job on Cloudflare covers the next layer.
If you would like a hand, our teams do Cloudflare development and AI-augmented development. That said, the three model pages and the bindings page above are enough to build this on your own.