AI & Creative Tools

We Dismissed This Model in September. It Is the One That Actually Fits Your Installation

laya is a 421M-parameter non-autoregressive decision model that answers in 33ms on a 2018 GPU, across 100+ languages — and it predates every model in the category we have been covering.

Our tracker has been showing convaiinnovations/laya climbing the likes table for two and a half weeks. Today its downloads moved too — +74%, 11,733 to 20,386 — and it has 5,275 likes, which is more than Cloudflare’s Clef (1,522), JEV-27B-VL (883) and GEV-26B-Decide (471) combined.

So we finally read the model card, and have to begin with a correction.

The correction

In our round notes for 21 September we recorded this:

convaiinnovations/laya — biggest likes gain in the tracker (+937). Created 2026-09-18, Apache-2.0, but it’s a text-classification / routing / guardrails model. Not creative-relevant.

That was wrong. Nine days later we wrote up Clef as the arrival of an interesting new category — models that return typed answers with calibrated probabilities instead of prose — and then three more pieces developing it. In every one of them we noted the same gap: these are 26–27B models, they need a real GPU, and for installation hardware something much smaller was still needed.

laya is that model, it was already there, and we had logged it and walked past it. The error was reading the tags — classification, routing, guardrails — as a product category rather than as a description of what typed decisions are for.

What it actually is

A multilingual, non-autoregressive System 1 decision model. It takes text plus typed questions and returns typed answers with calibrated probabilities in a single forward pass, in ~33ms, across 100+ languages.

And: it never produces text, which the card notes eliminates hallucination risk by construction.

Three checkpoints:

BackboneParamsContext
laya (root)ModernBERT-large421M512 tokens, English-optimised
laya-multilingualmmBERT-base322M1,024 → 8,192, 100+ languages, ~2.2× faster
laya-typed-decisionsModernBERT-large421Mspecialised for four workflows

Architecture: option markers at [MASK] tokens for scoring, a 2-layer transformer decision head, and an act/escalate head.

Apache 2.0. pip install laya.

The numbers that matter

32.8ms on a Tesla T4. A T4 is a 2018 inference card with 16GB — the cheapest thing on every cloud provider and the sort of GPU that turns up in a donated workstation. The card reports TypeSafe Jev at 236–276ms on the same hardware, making laya roughly 7.8× faster.

Accuracy 0.766 for the fine-tuned laya-typed-decisions on 2,000 typed-decision cases — above its own teacher ceiling of 0.735, and above Jev’s published 0.727.

Expected Calibration Error 0.081 after temperature scaling.

That last figure is the one to care about, and it is the one every other model in this category has asserted rather than measured. ECE is the gap between stated confidence and observed accuracy — bucket the predictions by claimed probability, compare each bucket’s actual hit rate, average the discrepancy. 0.081 means that when laya says 80%, reality is within about eight points of that. Every Clef/JEV/GEV piece we have written ended with check the calibration on your own data because nobody has published a number. Here is a published number.

RLCD, which is why the calibration is real

Reinforcement Learning for Calibrated Decisions trains the policy with zero-mean Gaussian noise exploration and strictly proper scoring rules — log, spherical, and ranked probability score — using REINFORCE with a group-mean baseline.

The phrase doing the work is strictly proper scoring rule, and it is a precise mathematical idea rather than a marketing one.

A scoring rule grades a probabilistic forecast. It is proper if the forecaster maximises their expected score by reporting their true belief, and strictly proper if that is the only way to maximise it. The log score and the Brier score are strictly proper; raw accuracy is not.

Why that matters: if you train a model on accuracy, the optimal strategy is to output 1.0 for its best guess every time, because partial credit for honest uncertainty does not exist. Confidence becomes meaningless. Train on a strictly proper rule and overconfidence is punished — saying 0.99 and being wrong costs far more than saying 0.6 and being wrong.

So calibration is not a post-hoc fix here; it is what the reward function selects for. That is a genuinely better design than quantising a large generative model and hoping its logits mean something.

The honest limitations, which the card states

Base checkpoints score near chance on typed decisions — 0.362 — and require fine-tuning for domain specialisation. This is the big one. laya out of the box is not a drop-in Clef replacement; it is a small, fast, well-calibrated substrate you fine-tune for your task.

High-cardinality choices underperform. Above about 77 options, token budgets bite, and the card concedes that Jev is better at 50+ option label spaces without tuning.

So the trade is clear: laya if your decision is between a handful of options and you can fine-tune; a 27B model if you need many options and zero tuning.

Why this is the one for installation work

Everything we have written about this category has carried the same caveat about hardware. laya removes it.

421M parameters is a few hundred megabytes. It runs on a laptop, on a Jetson, plausibly on a Raspberry Pi 5 with patience, and comfortably on any GPU made in the last decade.

33ms is inside a frame at 30fps, with room to spare.

100+ languages matters for public work. An installation in a museum with international visitors asking a system to make a semantic judgement about text is a real case, and most small models are English-only.

And “never produces text” is a safety property, not a limitation. An unattended piece that cannot emit a sentence cannot emit an embarrassing one. For anything in a public venue, that is worth more than capability.

The practical path: take laya-multilingual, fine-tune on a few hundred examples of the judgement your piece needs to make, and you have a calibrated classifier with an escalate head running locally at frame rate. That is a better answer than anything we have recommended in the last two weeks, and we should have found it a fortnight ago.