Our daily tracker picked up a new entrant on 8 October: autotrust/GLM5.3-Flash-E224-DGX-Spark, created 6 October 2026, already at 5,422 downloads and 242 likes. It is an unofficial derivative of GLM-5.3-Flash, and the interesting thing about it is that it was not designed around a parameter count. It was designed around a specific piece of hardware sitting on a desk.
The construction
GLM-5.3-Flash is a mixture-of-experts model with 288 routed experts per layer. This derivative keeps 224 of them — the E224 in the name — and discards 64 per layer, with the selection made by neural architecture search rather than by a magnitude heuristic.
That single change takes the model from 306 GiB to 141.5 GiB. Two further decisions do the rest:
- NVFP4 for the routed experts. A 4-bit floating point format with 16-element groups, native to Blackwell tensor cores. The group size is the point — it is small enough that outliers inside a block do not destroy the scale for everything around them, which is the usual failure mode of aggressive weight quantisation.
- BF16 for attention, the shared experts, and the vision tower. Precision is kept exactly where it is cheap in bytes and expensive to lose.
Active parameters stay at 18B via top-8 expert selection, unchanged from the original. Pruning the pool does not change how many experts fire per token — it changes how many are available to be chosen. The published benchmarks are described as within noise of the unpruned model: GPQA-Diamond around 91%, AIME 2025 at 88.3%. Licence is MIT.
The hardware target is the design brief
This is the part worth sitting with. The model is built to run on:
- two DGX Sparks linked over ConnectX-7 — 128 GB unified memory each, 256 GB combined, leaving roughly 70 GiB per node and about 40 GiB per node for KV cache; or
- a single B200 or GB200 with at least 180 GB.
And explicitly not on one DGX Spark. 141.5 GiB does not fit in 128 GB, and no amount of wanting it to will change that.
NVIDIA’s own walkthrough of linking two DGX Sparks — the configuration this model is sized for.
So the honest framing is not “a small local model.” This needs two machines that cost around the price of a car between them, and it is the opposite end of the spectrum from the 400M-parameter releases that have been driving the tracker for the last fortnight.
What makes it belong in the same conversation is the direction the engineering runs. The normal sequence is: train a model, then ask what hardware can serve it, then quantise until it fits something. Here the sequence is inverted — the memory envelope is the specification, and the architecture is cut to meet it. Neural architecture search is not being used to find a better model in the abstract; it is being used to find the best model that fits in 256 GB minus KV cache.
That inversion is the generalisable idea, and it scales down as well as up. The same method pointed at 24 GB, or 16, or 8, produces a differently-shaped answer from “take the big model and quantise harder” — because expert pruning removes capacity the model was not using for your workload, while uniform quantisation degrades everything the model does use.
What it is actually for
The stated target is interactive, low-concurrency desktop serving with adjustable thinking budgets. That phrasing deserves unpacking, because it describes a deployment pattern that barely existed two years ago.
Low concurrency means one person, maybe two. Not a service. The ~40 GiB of KV cache per node is a lot of context for one conversation and almost nothing for fifty, and the design clearly chooses the former.
Adjustable thinking budgets means the latency-versus-quality dial is in the operator’s hands rather than fixed by a provider’s cost model.
Together they describe a private, multimodal, reasoning-capable model — it is image-text-to-text, with the vision tower kept at BF16 — that answers to nobody’s rate limit. For creative work with material you cannot upload, that is the whole proposition, and the reason the cost of two Sparks is a comprehensible trade rather than an absurd one.
The caveats, stated plainly
It is unofficial. autotrust is not Zhipu AI. The pruning, the search, the quantisation and the benchmarks are all third-party work, and the benchmark figures are the publisher’s own.
“Within noise” is a claim about aggregate scores. GPQA-Diamond and AIME measure reasoning on a narrow slice of tasks. What 64 removed experts per layer cost on anything not in those benchmarks — rarer languages, unusual domains, long-tail factual recall, image understanding outside the common cases — is unmeasured. Pruning by architecture search optimises for the objective you give it, and capacity that served the tail is exactly the capacity a search will find expendable.
Expert-pruned derivatives are hard to compare. Tags on the repo include expert-pruned, nvfp4, modelopt, dgx-spark, gb10 and blackwell, which is an accurate description of a build with several simultaneous modifications — and no way to attribute a difference in output to any one of them.