AI & Creative Tools

A 78B Mixture-of-Experts That Runs 3.46B Parameters Per Token, in German and English

Aleph Alpha's Kolibri-1 landed on 3 October with configurable reasoning effort, a million-token extended context, and Apache 2.0 weights. Likes rose 274% in a day.

Our Hugging Face tracker put Aleph-Alpha/Kolibri-1 near the top of today’s likes table — 91 to 340 in a day, up 274%, and 407 by the time we checked the card directly. Released 3 October 2026, which makes it about a day old.

Aleph Alpha is the German lab that has spent several years as Europe’s most-discussed attempt at a sovereign alternative to the American labs. Kolibri-1 — Kolibri is German for hummingbird — is their mixture-of-experts reasoning model.

The numbers

Total parameters78 billion
Active per token3.46 billion
Layers50
Experts per layer384
Routed per token6, plus 1 shared
Native context262,144 tokens
Extended context1,048,576 tokens
LanguagesGerman and English
LicenceApache 2.0 — weights and config files only

Reported benchmarks: 96.5% average on English AIME (maths), 89.3% average on coding, with strong knowledge-task results, competing with larger dense models on fewer active parameters.

384 experts is an unusual number, and the ratio is the story

The headline MoE figure everyone quotes is total parameters. The number that determines what it costs you to run is active parameters, and 78B total / 3.46B active is a ratio of about 22:1 — very sparse by current standards.

Two consequences, and they pull in opposite directions:

Compute is cheap. 3.46B active means each token costs roughly what a 3.5B dense model costs to compute. That is small — comfortably within reach of a single consumer GPU’s compute budget.

Memory is not. All 78B parameters have to be resident, because the router can select any expert for any token. At 8-bit that is ~78GB before context; at 4-bit, ~39GB plus overhead. So this is not a model you run on a 24GB card without aggressive quantisation and offloading, and the fast-compute/large-memory profile is the defining practical characteristic of every sparse MoE.

384 experts per layer with 6 routed is finer-grained than the designs that popularised MoE (Mixtral used 8 experts, 2 routed). The current research direction favours many small experts over few large ones — more specialisation, better capacity utilisation — at the cost of a harder routing problem and more awkward memory access patterns. 384 is at the aggressive end of that trend.

The one shared expert is also a now-standard detail worth knowing: one expert every token passes through, which handles the general-purpose work so the routed experts can specialise rather than all having to learn the basics.

Configurable reasoning effort is the feature to care about

Kolibri-1 exposes low / medium / high reasoning effort levels built into the chat template, plus tool calling.

This matters more for practical work than any benchmark number. Reasoning models have one dominant failure mode in production: they think when you did not want them to. Ask a reasoning model to reformat a list and it may spend two thousand tokens considering the problem. For an interactive application — and certainly for anything in an installation loop — that is fatal, and the usual workaround is prompt engineering that begs the model to be brief.

A documented effort dial in the chat template means it is a parameter, not a negotiation. You set low for classification and formatting, high for the thing that actually needs thought, and your latency becomes predictable.

Read the licence carefully

Apache 2.0 applies to weights and configuration files only, with Aleph Alpha retaining rights to architecture and training methods.

That is a real distinction and worth being precise about. You can use, modify, fine-tune and deploy the weights commercially under Apache terms — which is the thing most people need, and it is genuinely permissive.

What is not given away is the architecture and training methodology as intellectual property. In practice this affects almost nobody who wants to use the model, and it does affect anyone who wanted to reimplement the approach and publish it. It is a narrower grant than “Apache 2.0” alone implies, and the card is upfront about it, which is better than the alternative.

Why German matters here

German and English rather than a long multilingual list is a deliberate scoping decision, and for European work it is the point.

Most open models are English-first with other languages as a long tail of varying quality. If you are building something for a German-speaking audience — an installation in Berlin with a text interface, a tool for German-language archives, anything where the output is read by native speakers — a model that treats German as a primary language rather than a translation target is materially different. The same argument applies to every non-English European language and is the actual case for sovereign model development, independent of any politics attached to it.

What it is and isn’t for

Aleph Alpha positions it for human-in-the-loop applications: conversational assistants, document processing, RAG, and agentic workflows where outputs receive human review before action.

That caveat is in their own card and should be respected. It is not positioned as something to put in an unattended loop making consequential decisions — which, notably, is exactly what the typed-decision models we wrote about yesterday are built for. Different tools, and the labs are being clear about which is which.

For creative work the realistic uses are the unglamorous ones: processing a corpus, summarising archives, driving a text-heavy interface in German, or as the reasoning layer behind a tool where a person reads the result. Running it locally needs real hardware; the Apache weights mean you can.