Our Hugging Face tracker put Cloudflare/clef at the top of today’s likes table with a move that is hard to miss: 19 likes to 632 in twenty-four hours, on 824 downloads. Its sibling Cloudflare/clef-flash entered the tracked set at trending rank #16. Both were created 30 September and modified 1 October.
Cloudflare has released a model that does not generate text.
Fahd Mirza runs Clef 27B locally — the useful look, since the whole point is that it is weights you can host.
What it does
Clef is a 27-billion-parameter multimodal model that converts a situation plus a question schema into probabilistic decisions. It accepts text, JSON, images or video, and returns a probability for each answer option — and nothing else. No free-form prose.
Architecturally it is Qwen/Qwen3.8-27B with an added joint schema head: a transformer component that routes evidence to questions and scores options jointly rather than one at a time.
Apache 2.0, matching the base model. That is the detail that makes the rest of it matter.
Why “typed output” is a different product, not a formatting option
You can already make a chat model emit structured output. Constrained decoding, JSON schema modes, function calling — every serious inference stack has some version of it.
But all of those are a text generator wearing a seatbelt. The model is still predicting tokens; the grammar is preventing it from predicting invalid ones. Two consequences follow, and both are familiar to anyone who has shipped this:
You get a choice, not a distribution. The model picks "category": "B" and gives you no honest signal about whether B beat A by a mile or by a hair. You can scrape token logprobs and squint, but they are the probabilities of tokens, not of answers, and they are badly calibrated for the purpose.
And it is slow, because it is still writing. Generating a structured response means producing every token of it autoregressively.
Clef’s joint schema head addresses both. Scoring options jointly means the options are compared against each other in one pass — which is what you want, because “is this A or B” is a comparison, not two independent judgements. And you get a probability per option, which is a number you can threshold, route on, abstain on, or log.
The latency number is the headline
| Median latency | |
|---|---|
| Clef | 209.3 ms |
| Clef-Flash | 38.8 ms |
Cloudflare reports Clef-Flash as comparable in accuracy on many benchmarks, with Clef sometimes better on specialised tasks.
38.8 milliseconds changes what you can build. That is inside a video frame at 24fps. It means a model decision can sit inside an interactive loop rather than beside it — you can ask a question of a camera feed, a sensor reading or a user action and act on the answer before the next frame.
Performance is reported on a Decision Index benchmark suite spanning 45+ datasets, with strongest results in function-calling, knowledge QA and reasoning.
Why a creative-technology site cares about a business-workflow model
Cloudflare built this for routing, classification and automated decisions — their own problem, at their own scale. The capability is more general than the pitch.
Installations are decision machines. Almost every interactive piece is a loop of look at the situation → decide what to do → do it. The decision step is usually a pile of hand-tuned thresholds on sensor values, because the alternative — a language model — was too slow, too chatty, and too unpredictable to put in the loop.
A model that takes an image plus a question schema and returns calibrated probabilities in 38 milliseconds is a different proposition. Concretely:
- Scene and situation classification that is semantic rather than pixel-based. Not “is there motion in region 3” but “is someone waiting, passing through, or leaving” — a distinction no blob detector can make.
- Routing between behaviours in a piece with several modes, with a probability you can hysteresis on so it does not flap between states.
- Abstention. This is the underrated one. A probability distribution lets the work say I don’t know and fall back to a default, which is the single most useful thing an unattended installation can do. A chat model asked to classify will always confidently classify.
- Video input, which the card lists, meaning temporal questions are in scope rather than frame-by-frame guessing.
We have covered two other points on this curve recently — SupersonicLabs’ Julia-1 at 144M parameters, built only for typed decisions, and Mapika’s decider-4b. Clef is the same idea with a large company’s resources behind it, and the three together look like a category forming rather than a one-off.
The caveats
27B is not small. Running Clef locally means a serious GPU, or quantisation and the quality questions that come with it. Clef-Flash is “smaller and faster” but the card does not state its parameter count in the material we have. For an installation on a modest machine, Julia-1’s 144M is still the more realistic answer; Clef is for when you have a real GPU or a server.
“Probabilistic” is not the same as “calibrated.” A model that outputs numbers summing to one is not thereby honest about its uncertainty. Before you build abstention logic on Clef’s probabilities, check them against your own data — plot predicted confidence against actual accuracy and see whether 0.8 means 0.8.
And the benchmark is Cloudflare’s. A 45-dataset suite called the Decision Index, published by the people who built the model that tops it, is a reasonable starting point and not an independent result.
Those noted: an Apache 2.0, multimodal, probability-returning decision model with a sub-40ms variant is a genuinely useful thing to exist, and the licence means you can find out for yourself.