Cloudflare cut the price of Clef-flash, now its lowest-priced decision model, from $0.09 to $0.038 per million input tokens on 9 October 2026, and paid for part of the cut by shrinking the hosted model’s context window from 64k to 24k tokens. The same post launched Clef-omni, a decision model that takes audio and video as well as text and images.
Cloudflare’s post says the new price makes Clef-flash “cheaper than Jev”, the decision model from TypeSafe. Decision models return scored choices for an agent instead of generated text, and Cloudflare’s docs say Clef models do not charge for output tokens.
Business Pill · MODEL ROUTING
A short explainer of model routing: sending each task to the model that fits it on cost and capability. It relates to Cloudflare’s advice that long inputs go to Clef rather than Clef-flash. It teaches the general idea only and says nothing about any company in this story.
The key insight: As we read it, Cloudflare is pricing the decision step of an agent close to zero and charging for the long-context and multimodal versions. The post ties the cheaper Clef-flash price directly to a smaller hosted context window.
What Cloudflare Changed
The post lists three prices per million input tokens. Clef-flash went from $0.09 to $0.038. Clef stays at $0.24. Clef-omni launched at $0.15.
The Workers AI docs pages show the same figures: $0.038 per M input tokens for Clef-flash, with a 24,576-token context window, and $0.15 per M input tokens for Clef-omni, with a 64,000-token window.
The post calls the context cut a trade-off: “one trade-off we have to make in order to have cheaper pricing is to cut the context window of Clef-flash.” The hosted Clef-flash now has a 24k window, “instead of 64k as previously advertised.”
Cloudflare says the change touches few requests: “only 0.24% of requests exceed 24k input tokens.” The weights on Hugging Face are unchanged and, per the post, were trained to support a 256k window for anyone who self-hosts.


Clef-omni Adds Audio and Video
Clef-omni accepts audio (wav or mp3) and video (mp4 or webm) alongside text and images in one call, the post says. Cloudflare built it on Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts model, and released the weights on Hugging Face.
The post gives its speeds: text decisions in about 130 ms at the median, images in about 150 ms, audio clips in a few hundred milliseconds, and a 21-second video clip with sound in about 1.5 seconds.
The docs say media is billed as input tokens: audio at about 780 tokens a minute, and video at up to about 15,400 tokens a minute at maximum resolution.
A Faster Clef
Cloudflare says it made the larger Clef model faster without new weights, mostly through changes to the serving layer, including a move to SGLang.
Its table shows median latency falling from 262 ms to 152 ms at about 800 tokens, from 616 ms to 305 ms at about 3,400 tokens, and from 2,721 ms to 1,635 ms at about 16,000 tokens.
How the Prices Compare
Microsoft’s Decision-1 post of 9 October prices input tokens at $0.042 per million, with output tokens free. By our arithmetic, Clef-flash’s new $0.038 is $0.004 below that.
Cloudflare’s 1 October launch post released Clef and Clef-flash under an Apache 2.0 license and described them as “fully Jev-API compatible”. It gave Clef a 64k context window, “compared to Jev’s 32k”.
Cloudflare’s own benchmark table on TypeSafe’s workflow evals shows mixed results: on invoice processing, Clef scores 64.7 exact actions against Jev’s 61.8, while on agent trace observability Jev scores 71.6 against Clef-omni’s 65.8.
The Structural Read
The post presents the context cut as the price of the price cut. Cloudflare says only 0.24% of requests exceed 24k input tokens, and it points users with longer inputs to Clef, which keeps its 64k window at $0.24.
The family now spans three price points per million input tokens: $0.038 for Clef-flash, $0.15 for Clef-omni and $0.24 for Clef. The docs say Clef models do not charge for output tokens.
The weights stay open. The post says the Hugging Face weights for Clef-flash are untouched and trained for a 256k window, so the 24k limit applies to Cloudflare’s hosted version only.
Cloudflare blog, 9 October 2026
“one trade-off we have to make in order to have cheaper pricing is to cut the context window of Clef-flash”
Three Implications
FOR TEAMS BUILDING AGENTS The docs price Clef models on input tokens only, so the cost of a decision depends on how much context and media goes into the call.
FOR LONG INPUTS Cloudflare says requests over 24k input tokens should move from Clef-flash to Clef, which keeps a 64k window at $0.24 per million input tokens.
WHAT TO WATCH The 9 October post does not state Jev’s price, so how the “cheaper than Jev” claim holds will depend on TypeSafe’s published pricing.
The Business Engineer Lens
This story maps onto the Business Engineer framework The Open vs Closed Meta-Framework.
The framework’s starting point: “Close the SCARCE layer (keep proprietary) + Open the ABUNDANT layer (commoditize)”.
As we read it, Cloudflare opens the model weights under Apache 2.0 and charges for the hosted service, where it sets the context window and the price.
What Is Not Established
We read Cloudflare’s 9 October post and its 1 October launch post in full, the Workers AI docs pages for Clef-omni and Clef-flash, and the text of Microsoft’s Decision-1 post. The benchmark figures are Cloudflare’s own.
The 9 October post we read does not state Jev’s price, so the “cheaper than Jev” comparison is Cloudflare’s. The post does not give usage numbers for the Clef models beyond the 0.24% share of requests above 24k tokens. We did not contact Cloudflare or TypeSafe.
The Bottom Line
Cloudflare cut Clef-flash to $0.038 per million input tokens and narrowed its hosted context window from 64k to 24k tokens, saying only 0.24% of requests exceed 24k. It launched Clef-omni at $0.15 with audio and video input, and made Clef faster. By our arithmetic, Clef-flash is now $0.004 below the price Microsoft gives for Decision-1.
94,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
A note on sourcing. We read Cloudflare’s posts of 1 and 9 October 2026, the Workers AI docs pages for Clef-omni and Clef-flash, and the text of Microsoft’s Decision-1 post; the prices, speeds and benchmark scores are the companies’ own. We did not use press coverage, and we did not contact Cloudflare or TypeSafe. Nothing here is a forecast, and nothing here is financial or investment advice.
Sources: Cloudflare: Introducing Clef-omni, a faster Clef and a cheaper Clef-flash (9 Oct 2026) · Workers AI docs: Clef-flash · Workers AI docs: Clef-omni · Cloudflare: Introducing Clef (1 Oct 2026) · Microsoft: Microsoft-Decision-1 (9 Oct 2026)








