AI & Creative Tools

Two 27B Decision Models Landed on the Same Day, on the Same Base, and Now Argue Over the Same Benchmark

AutoTrust's JEV-27B-VL and Cloudflare's Clef both appeared on 30 September, both on Qwen3.8-27B, both Apache 2.0. The scoreboard they fight on belongs to one of them.

Yesterday we wrote up Cloudflare’s Clef — a 27B model that returns a probability per answer option and no prose — on the strength of the largest single-day likes move our tracker has recorded. Today the tracker has a follow-up that reframes it.

autotrust/JEV-27B-VL downloads +72%, 179,237 → 308,531. And both clef and clef-flash entered Ollama’s popular list, at ranks #6 and #7, two days after release — which means the weights are now a one-line pull rather than a Hugging Face download.

First, a correction

Yesterday we described the Decision Index as “a 45-dataset suite called the Decision Index, published by the people who built the model that tops it,” and treated that as a reason for caution.

That was wrong, and wrong in the direction that matters. The Jev Decision Index was created by Typesafe AI — the makers of Jev, Clef’s direct competitor. Cloudflare is reporting its results against a rival’s benchmark, which is a considerably stronger position than publishing your own. Cloudflare’s blog post states plainly that Clef “is currently the leader when evaluated against” it, and links the full results.

Reporting against the incumbent’s yardstick is the good version of this. Apologies for getting it backwards.

What the scoreboard actually says

Cloudflare’s own figures, which are worth reading closely because the pattern is more interesting than the headline:

BenchmarkResult
BANKING77 (macro-F1)Clef 94.20 vs Jev 79.74
CLINC150+OOS (macro-F1)Clef 97.43 vs Jev 89.27
API-Bank (accuracy)Clef-flash 93.1%, Clef 91.9%, Jev 88.2%
When2Call (accuracy)Jev 80.97 vs Clef 72.37, Clef-flash 65.58

Cloudflare reports winning or tying 6 of 10 in its main table, and losing on When2Call, BRIGHT and PhishNChips. On four end-to-end business workflows it leads invoice processing (64.7 vs 61.8) and security incidents (62.9 vs 61.7), while Jev keeps agent trace observability at 71.6.

Latency, across 43 benchmarks: Clef-flash 38.8ms median (122.4ms at p95), Clef 209.3ms, Jev 524.1ms.

The When2Call loss is the informative one. Classification benchmarks — BANKING77, CLINC150 — are where Clef wins by wide margins. When2Call asks a harder question: should a tool be called at all, or should the model decline? Jev wins that by eight points. So the picture is not “Clef is better”; it is Clef is better at deciding between options and Jev is better at deciding whether to act. Those are genuinely different capabilities and the second is the one that keeps a system from doing something stupid.

Also worth correcting from yesterday: Clef-flash is built on Qwen3.5-9B, not an unspecified smaller variant. Clef is on Qwen3.8-27B.

The category, which is the actual story

JEV-27B-VL is from AutoTrust AI Lab, built on their earlier text-only JEV-27B. Structurally it is strikingly close to Clef:

  • System 1 — fast typed decisions: yes/no, multiple choice from 2–256 options, or 0–5 ratings, returning calibrated probabilities in one forward pass
  • System 2 — reasoning responses from the underlying Qwen3.8-27B
  • Apache 2.0, same base, same licence

Benchmarks on its card: 78.3% on VL-RewardBench multimodal judging, 73.2% on Plan-RewardBench agent evaluation, 95% on computer-use tasks, 75% on robotic pick-and-place.

So within a week we have Clef, Clef-flash, JEV-27B-VL, the earlier text-only JEV, and reporting elsewhere of an Amazon Strands Decider in the same space. Add SupersonicLabs’ Julia-1 at 144M and Mapika’s decider-4b, both of which have passed through this tracker in the last fortnight, and the shape is unmistakable: a model category formed in about four weeks.

Two things are notable about how it formed. Everyone picked the same base model — Qwen3.8-27B, Apache 2.0 — which tells you the differentiation is entirely in the decision head and the post-training, not the backbone. And a benchmark emerged before the category settled, with the first mover’s suite becoming the shared yardstick, which is historically how these things go and is why Typesafe’s position is stronger than its win-rate suggests.

What this is worth to someone making things

The practical answer has improved a lot in two days, because of Ollama.

ollama pull clef-flash is a different proposition from a 27B Hugging Face download and a GPU. A 9B model answering in 38.8ms, pullable in one command, that takes an image and a question schema and returns calibrated probabilities over up to 256 options, is now genuinely available for the thing interactive work needs: a semantic decision inside a frame budget.

The honest caveats stand from yesterday and one is now sharper:

Calibration still needs checking on your own data. Both cards use the word “calibrated”; neither is a guarantee about your inputs. If you intend to build abstention logic — fall back to the default when confidence is below 0.7 — plot predicted confidence against observed accuracy before trusting the threshold.

And the benchmark fragility problem is real. We are publishing a piece today about a paper showing that in text-to-3D evaluation, how you measure varies more than which system you use for 17 of 19 evaluators. There is no reason to think decision-model benchmarks are immune. A 6-of-10 win on one suite, however independently constructed, is a reason to test rather than a reason to conclude.