Our Hugging Face tracker put Aleph-Alpha/Kolibri-1 near the top of today’s likes table — 91 to 340 in a day, up 274%, and 407 by the time we checked the card directly. Released 3 October 2026, which makes it about a day old.
Aleph Alpha is the German lab that has spent several years as Europe’s most-discussed attempt at a sovereign alternative to the American labs. Kolibri-1 — Kolibri is German for hummingbird — is their mixture-of-experts reasoning model.
The numbers
| Total parameters | 78 billion |
| Active per token | 3.46 billion |
| Layers | 50 |
| Experts per layer | 384 |
| Routed per token | 6, plus 1 shared |
| Native context | 262,144 tokens |
| Extended context | 1,048,576 tokens |
| Languages | German and English |
| Licence | Apache 2.0 — weights and config files only |
Reported benchmarks: 96.5% average on English AIME (maths), 89.3% average on coding, with strong knowledge-task results, competing with larger dense models on fewer active parameters.
384 experts is an unusual number, and the ratio is the story
The headline MoE figure everyone quotes is total parameters. The number that determines what it costs you to run is active parameters, and 78B total / 3.46B active is a ratio of about 22:1 — very sparse by current standards.
Two consequences, and they pull in opposite directions:
Compute is cheap. 3.46B active means each token costs roughly what a 3.5B dense model costs to compute. That is small — comfortably within reach of a single consumer GPU’s compute budget.
Memory is not. All 78B parameters have to be resident, because the router can select any expert for any token. At 8-bit that is ~78GB before context; at 4-bit, ~39GB plus overhead. So this is not a model you run on a 24GB card without aggressive quantisation and offloading, and the fast-compute/large-memory profile is the defining practical characteristic of every sparse MoE.
384 experts per layer with 6 routed is finer-grained than the designs that popularised MoE (Mixtral used 8 experts, 2 routed). The current research direction favours many small experts over few large ones — more specialisation, better capacity utilisation — at the cost of a harder routing problem and more awkward memory access patterns. 384 is at the aggressive end of that trend.
The one shared expert is also a now-standard detail worth knowing: one expert every token passes through, which handles the general-purpose work so the routed experts can specialise rather than all having to learn the basics.
Configurable reasoning effort is the feature to care about
Kolibri-1 exposes low / medium / high reasoning effort levels built into the chat template, plus tool calling.
This matters more for practical work than any benchmark number. Reasoning models have one dominant failure mode in production: they think when you did not want them to. Ask a reasoning model to reformat a list and it may spend two thousand tokens considering the problem. For an interactive application — and certainly for anything in an installation loop — that is fatal, and the usual workaround is prompt engineering that begs the model to be brief.
A documented effort dial in the chat template means it is a parameter, not a negotiation. You set low for classification and formatting, high for the thing that actually needs thought, and your latency becomes predictable.
Read the licence carefully
Apache 2.0 applies to weights and configuration files only, with Aleph Alpha retaining rights to architecture and training methods.
That is a real distinction and worth being precise about. You can use, modify, fine-tune and deploy the weights commercially under Apache terms — which is the thing most people need, and it is genuinely permissive.
What is not given away is the architecture and training methodology as intellectual property. In practice this affects almost nobody who wants to use the model, and it does affect anyone who wanted to reimplement the approach and publish it. It is a narrower grant than “Apache 2.0” alone implies, and the card is upfront about it, which is better than the alternative.
Why German matters here
German and English rather than a long multilingual list is a deliberate scoping decision, and for European work it is the point.
Most open models are English-first with other languages as a long tail of varying quality. If you are building something for a German-speaking audience — an installation in Berlin with a text interface, a tool for German-language archives, anything where the output is read by native speakers — a model that treats German as a primary language rather than a translation target is materially different. The same argument applies to every non-English European language and is the actual case for sovereign model development, independent of any politics attached to it.
What it is and isn’t for
Aleph Alpha positions it for human-in-the-loop applications: conversational assistants, document processing, RAG, and agentic workflows where outputs receive human review before action.
That caveat is in their own card and should be respected. It is not positioned as something to put in an unattended loop making consequential decisions — which, notably, is exactly what the typed-decision models we wrote about yesterday are built for. Different tools, and the labs are being clear about which is which.
For creative work the realistic uses are the unglamorous ones: processing a corpus, summarising archives, driving a text-heavy interface in German, or as the reasoning layer behind a tool where a person reads the result. Running it locally needs real hardware; the Apache weights mean you can.