A model that answers questions with probabilities instead of prose is a genuinely useful thing to have in an installation. It can classify a situation, route between behaviours, score an input, or decline to decide — and it cannot produce an embarrassing sentence, because it cannot produce sentences.
The problem with the ones that have arrived recently is size. Cloudflare’s Clef is 27B. JEV-27B-VL is 27B. GEV-26B-Decide is 26B with ~4B active. All of them need the full weights resident on a real GPU, which rules out most of the hardware this work actually runs on.
convaiinnovations/laya is 421 million parameters. It is the one that fits.
What it is
A multilingual, non-autoregressive System 1 decision model from Convai Innovations. It takes text plus typed questions and returns typed answers with calibrated probabilities in a single forward pass, in ~33ms, across 100+ languages.
It never produces text, which eliminates hallucination by construction rather than by guardrail.
Three checkpoints:
| Backbone | Params | Context | |
|---|---|---|---|
laya (root) | ModernBERT-large | 421M | 512 tokens, English-optimised |
laya-multilingual | mmBERT-base | 322M | 1,024 → 8,192, 100+ languages, ~2.2× faster |
laya-typed-decisions | ModernBERT-large | 421M | specialised for four workflows |
Architecturally: option markers at [MASK] tokens for scoring, a 2-layer transformer decision head, and an act/escalate head.
Apache 2.0. pip install laya.
The numbers
32.8ms on a Tesla T4. A T4 is a 2018 inference card — the cheapest option at every cloud provider, and the kind of GPU that turns up in a donated workstation. The model card reports TypeSafe Jev at 236–276ms on the same hardware, making laya roughly 7.8× faster.
Accuracy 0.766 for the fine-tuned laya-typed-decisions on 2,000 typed-decision cases — above its own teacher ceiling of 0.735, and above Jev’s published 0.727.
Expected Calibration Error 0.081 after temperature scaling.
That last figure is the one that matters most, and it is the only measured calibration number anyone in this category has published. ECE is the gap between stated confidence and observed accuracy: bucket the predictions by claimed probability, compare each bucket’s actual hit rate, average the discrepancy. 0.081 means that when laya says 80%, reality is within about eight points of that.
Every other model here uses the word “calibrated” as a description. This one gives you an error bar on it.
RLCD, and why the calibration is structural
Reinforcement Learning for Calibrated Decisions trains the policy with zero-mean Gaussian noise exploration and strictly proper scoring rules — log, spherical, and ranked probability score — using REINFORCE with a group-mean baseline.
The load-bearing phrase is strictly proper scoring rule, and it is a precise mathematical idea rather than a marketing one.
A scoring rule grades a probabilistic forecast. It is proper if the forecaster maximises their expected score by reporting their true belief, and strictly proper if that is the only way to maximise it. The log score and the Brier score are strictly proper. Raw accuracy is not.
Why that distinction decides everything: train a model on accuracy and the optimal strategy is to output 1.0 for its best guess, every time, because there is no partial credit for honest uncertainty. Confidence becomes decoration. Train on a strictly proper rule and overconfidence is punished — saying 0.99 and being wrong costs far more than saying 0.6 and being wrong.
So calibration here is not a post-hoc temperature fix applied to a model that was optimised for something else. It is what the reward function selects for, which is a better foundation than quantising a large generative model and hoping its logits mean something.
The limitations, which the card states plainly
Base checkpoints score near chance on typed decisions — 0.362 — and require fine-tuning for domain specialisation. This is the important one. laya out of the box is not a drop-in Clef replacement; it is a small, fast, well-calibrated substrate you fine-tune for your task.
High-cardinality choices underperform. Above roughly 77 options, token budgets bite, and the card concedes that Jev is better at 50+ option label spaces without tuning.
The trade is therefore clear: laya if your decision is between a handful of options and you can fine-tune; a 27B model if you need many options and zero tuning.
Why this is the one for installation work
421M parameters is a few hundred megabytes. It runs on a laptop, on a Jetson, plausibly on a Raspberry Pi 5 with patience, and comfortably on any GPU made in the last decade.
33ms is inside a frame at 30fps, with room to spare.
100+ languages matters for public work. An installation in a museum with international visitors, making a semantic judgement about text, is a real case — and most small models are English-only.
And “never produces text” is a safety property. An unattended piece that cannot emit a sentence cannot emit an embarrassing one. In a public venue that is worth more than capability.
The act/escalate head is the detail that completes it. It gives you a built-in signal for I should not decide this — which is the single most useful thing an unattended system can do, and the thing GEV’s confidence gate arrived at from the opposite direction.
The practical path: take laya-multilingual, fine-tune on a few hundred examples of the judgement your piece needs to make, and you have a calibrated classifier with an escalation signal running locally at frame rate.