Our daily tracker flagged a new entrant on 9 October with an unusual signature: 139 likes against 89 downloads. Attention far ahead of use, on a repo created the day before.
Phocinae/Phocinae-Largha-150M-v1 — 斑海豹, the spotted seal — is a 144.3M-parameter decision model. It is not a chat model and does not generate text. You hand it a state plus a list of typed questions, and it returns answers with confidence scores in one forward pass.
The architecture is the news
Every model in this category so far has been a decoder. laya is a 421M non-autoregressive decoder; Cloudflare’s Clef, JEV and GEV are all decoder-family. Phocinae is an encoder — built on mmBERT-small, a ModernBERT variant: hidden 384 × 22 layers × 6 heads, a 256k-vocab tokenizer, RoPE with sliding-window plus full attention, 8192-token base context, with decisions capped at 512 tokens and a 192-token head attention window.
Three question types, and the typing is what makes an encoder sufficient:
noul— yes/nochoice— pick one of Nscore— a 2–10 rating
Because the answer space is enumerable, there is nothing to generate. The model scores the options and you read the scores. usage comes back as {"input_tokens": 146, "output_tokens": 0} — hence the repo’s zero-output-tokens tag, which is a literal description rather than a marketing flourish.
288.6 MB of weights. 1.6 GB peak inference VRAM. Apache 2.0, with the runtime published separately as phocinae-server, a FastAPI service bound to 127.0.0.1 that loads the weights with a plain PyTorch forward.
The numbers
On the typed-decisions benchmark (400 cases / 2,000 decisions):
| Model | en accuracy | basis |
|---|---|---|
| Phocinae-Largha | 0.906 | specialist — fitted on this dataset’s train split |
| meraGPT | 0.768 | generalist, zero-shot |
| Laya | 0.766 | self-measured, native interface |
| JEV | 0.727 | generalist, zero-shot |
Chinese: 0.848, on machine-translated cases.
Latency: 21.0 ms p50 on an RTX 5090 in fp16. CPU single-thread is 1.64 s per case (one case being a state plus five questions in a single forward pass); CPU 8-thread batching gives 8–20 decisions/s.
And the figure the whole thing is built around: a confidence gate at τ=0.6 that escalates low-confidence decisions to a real LLM and keeps the rest. On the kept subset, accuracy rises 0.906 → 0.9936 while LLM calls fall 55.0%. Full routed accuracy across everything, escalations included, is 0.8135.
Read the asterisk on the headline number
The card itself flags it, and it matters: Phocinae is a specialist fitted on this dataset’s train split, while Laya, JEV and meraGPT were measured zero-shot. 0.906 against 0.766 is not a like-for-like comparison — it is a model trained on this distribution against models that had never seen it.
That does not make the number meaningless. If your decisions resemble the typed-decisions distribution, a small specialist genuinely is the better tool, and that is the realistic deployment case: you are gating rm -rf or picking a tool from a list of ten, not answering open questions. But the gap is partly a measure of fine-tuning, not of architecture, and the card is explicit about which is which.
Two disclosures that are unusual to see published
The benchmark it fails. On JevBench public-231 it scores 0.5455 (126/231) — below the 58.4% gate, and the card says so under its own heading, with the words “disclosed honestly.” It also notes the one subtask where it is perfect: tool_selection 12/12, on n=12.
The bug it hasn’t fixed. The card carries a Known limitation in tool routing section stating that Router.route_tool() confuses semantically similar tools.
A benchmark failure and an open defect, in the README, next to the numbers that favour it. Across this entire category — and we have now covered laya, Clef, JEV, GEV and an expert-pruned GLM derivative — this is the first card to publish a result that undercuts its own case.
The calibration figures tell their own story
The card gives ECE twice:
- shipped column: 0.2519 (en) / 0.1941 (zh)
- with the bundled calibration column: 0.0168 / 0.0152
A fifteenfold improvement, from temperatures listed to seven decimal places and applied at inference.
The reason to care is that 0.2519 is bad calibration. A model with ECE around 0.25 that reports 0.9 confidence is wrong far more often than one time in ten, and the entire escalate-routing argument depends on confidence meaning what it says — a τ=0.6 gate on miscalibrated scores routes the wrong cases. Shipping both numbers, rather than only the good one, lets you see that the calibration column is load-bearing rather than cosmetic. Do not disable it.
For comparison, laya published an ECE of 0.081 and remains the reference point in this category for a measured calibration figure on an uncorrected output.
Option-order invariance is a real and under-tested property
The card measures what happens when you reorder the options in a choice question. On CPU fp32: flip150 0.0200, flip400 0.0217, random-mean 0.0144, any 0.0283 — lower is better.
This is the kind of thing almost nobody measures and everybody should. A decision model whose answer depends on whether “deny” came before or after “allow” is not making a decision; it is exhibiting a positional artefact. The bias is well known in LLM multiple-choice evaluation and is usually handled by averaging over permutations at test time, which you cannot do in a 21 ms production gate. Publishing a flip-robustness number means the property was treated as a requirement rather than discovered later.
What it is for
The stated use is the repetitive interior decisions of an agent session: command approvals, tool selection, step checks, output screening. Each of those is currently a 500–4,000-token API call that returns, in effect, a single token of information.
For creative work the relevant properties are the ones that follow from running locally: deterministic inference, so decisions are auditable and replayable, and nothing leaving the machine. If you are running an agent over client material, a local gate on 288 MB of weights is a different risk posture from a round trip to someone else’s endpoint — and at 21 ms against 1.5 s, it is also the faster one.
The honest caveats: it is en-first with Chinese evaluated largely on machine translation; the training mix is roughly 2,400 machine-translated Chinese rows against 1,400 native; CPU-only deployment at 1.64 s per case is not an interactive gate; and the tool-routing defect is unresolved.