Cloudflare ships Clef as open decision models
Clef decision models are Cloudflare's first in-house pair for yes-no, multiple-choice, and ranking questions, released October 1 on Workers AI and Hugging Face.

Clef decision models arrived on October 1 as Cloudflare's attempt to sit in the same lane as TypeSafe's Jev: a system that returns probabilities over a schema instead of writing a paragraph. The company blog, bylined by Michelle Chen, Alex Reneau, and Kevin Flansburg, says two models, Clef and Clef-flash, are on Workers AI, with weights on Hugging Face under Apache 2.0, plus a hands-on fine-tuning offer that is supposed to become a self-serve reinforcement-learning product later.
TL;DR - Cloudflare says the models use frozen Qwen backbones, a 64,000-token window, and a vision encoder Jev does not have. - The Register confirmed with Chen that training data is not public, and that local runs need about 41 GB or 85 GB of VRAM. - Cloudflare's own tables show wins on several tests and losses to Jev on When2Call and agent-trace observability.
What actually happened
The launch post is Introducing Clef. Cloudflare says these Clef decision models are built on Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash. Inference is a prefill-only pass, then a non-autoregressive score over the allowed choices. The company says that is why it can be faster than an ordinary chat model that has to emit tokens. In one internal example, a domain classification through Browser Rendering took 2.2 seconds on Clef versus 4.7 seconds on gpt-oss-120b, which also returned fewer labels.
Cloudflare's benchmark table is self-reported. On BFCL case-exact, it lists Clef at 98.47 and Clef-flash at 98.76 against Jev at 95.75. On BANKING77 macro-F1, Clef is 94.20 and Jev is 79.74. On CLINC150+OOS, Clef is 97.43 and Clef-flash falls to 66.77, under Jev's 89.27. On When2Call accuracy, Jev leads at 80.97 versus Clef at 72.37 and Clef-flash at 65.58. A second table of TypeSafe-style workflows has Clef ahead on invoice processing (64.7 versus 61.8) and slightly ahead on security incidents at 62.9 versus Jev at 61.7. Jev leads agent-trace observability at 71.6 versus Clef at 68.5, with Clef-flash at 69.8 on that same row. Median latency in Cloudflare's table is 209.3 ms for Clef, 38.8 ms for Clef-flash, and 524.1 ms for Jev. Those numbers should be read as the vendor's lab sheet. The Register says they had not been reproduced on the official Decision Index at publication.
Where the two pages split
Context length is a place where the two opened pages do not use the same sentence. Cloudflare says its window is 64,000 tokens "compared to Jev's 32k." The Register says Jev can also handle up to 64,000 tokens across a request, while Jev's state plus the longest individual question is limited to 32,000. Both descriptions are left standing. They are not averaged into a single context figure.
The Register's October 1 piece, published 21:39 UTC (12:39 a.m. MSK on October 2), is the independent account. Brandon Vigliarolo writes that Chen confirmed the training sets are not public, even though the weights are called open source in the announcement. Chen also gave hardware floors: Clef-flash on a GPU with at least 41 GB of VRAM, and Clef at 85 GB, assuming one request at a time and a 64,000-token window. The Register puts hosted Clef at $0.24 per million tokens, against Jev at $0.042 per million. That price gap is the Register's comparison, not a Cloudflare table in the post opened here.
The API example in the blog is a support ticket with three typed questions: urgent or not, a team choice, and a severity score. Cloudflare says the contract is Jev-compatible, so a caller can swap the endpoint without redesigning the question schema. It also says Workers AI will not read, store, or train on requests unless the customer opts into fine-tuning. The fine-tune path is not a button yet. The post describes a forward-deployed engineer engagement first, using AI Gateway logs, Workers AI rollouts, Containers as a sandbox, a new trainer, and bring-your-own-model hosting.
Readers comparing agent products can use the Robinhood Agents note and the MAI voice models note as separate examples of packaged AI features. Neither is a decision model, and neither is a benchmark for Clef.
What to do as a reader
If you are evaluating Clef decision models, separate three artifacts: the weight files, the hosted price, and the score table. The weights are the part Cloudflare actually put on Hugging Face. The $0.24 figure is reported by The Register. The score table is Cloudflare marking its own homework until an outside harness reruns it. A loss to Jev on When2Call and on agent-trace observability is in Cloudflare's own grid, so a "faster and smarter" headline does not cover every row.
Local hardware is the other filter. A 41 GB or 85 GB VRAM floor, on Chen's account to The Register, means a laptop demo is not the default. Hosted use adds a token bill that the Register says is several times Jev's posted input price. None of that is a reason to buy or sell a Cloudflare share. It is a reason to read the model card, the license, and the row you actually care about before treating a launch chart as a ranking.
Not financial advice. DYOR, ser.