AI & Creative Tools

Qwen-Image-2.1 Now Runs on an 8GB Card, and the Quantisers Got There Before the Ecosystem Did

Three days after the weights dropped, community GGUF quants pushed a 14GB model onto consumer GPUs. Our tracker shows the quantised repos outgrowing the official one by a wide margin.

We covered Qwen-Image-2.1 on September 20, when Alibaba open-sourced it: a 7-billion-parameter unified generation-and-editing model with native RGBA output. The story then was the transparency feature. Three days later there’s a second story, and it’s about distribution.

The original BF16 diffusion model is roughly 14.2GB. That is a 24GB-card proposition, comfortably outside what most people have in a desktop. It is now running on 8GB.

What the download numbers actually show

Our Hugging Face tracker compared September 22 to September 23. The quantised repositories are not trailing the official release — they are outpacing it:

RepositoryDownloads (22nd → 23rd)Change
Comfy-Org/Qwen-Image-2.11,429,925 → 2,220,609+790,684 (+55.3%)
abenzerps/Qwen-Image-2.1-Uncensored-GGUF182,313 → 350,678+168,365 (+92.3%)
pottokao/Qwen-Image-2.1-Text-Encoder-Heretic-GGUF34,059 → 74,291+40,232 (+118.1%)
leejet/Qwen-Image-2.1-GGUF22,693 → 44,260+21,567 (+95.0%)
Qwen/Qwen-Image-2.1 (official)16,242 → 28,407+12,165 (+74.9%)

The official repo gained twelve thousand downloads in a day. The ComfyUI-packaged version gained seven hundred and ninety thousand. That ratio is the whole point: almost nobody is downloading the reference weights. They are downloading the version that drops into the tool they already run.

How the quantisations work

The GGUF format came out of the llama.cpp world and arrived in image generation through the ComfyUI-GGUF custom node, which supplies a GGUF diffusion-model loader. You keep the standard Qwen-Image-2.1 text encoder, VAE, conditioning and sampler, and swap only the UNet loader.

Community releases now cover Q2 through Q8. The consensus recommendation is Q4_K_M as the size/quality sweet spot. At least one release ships a mixed-precision scheme specific to Qwen-Image-2.1, holding the more sensitive layers at higher precision so that aggressive quantisation doesn’t cost the text rendering — which is the capability Qwen-Image is most known for, and the first thing to break when you crush a diffusion model too hard.

The part worth being careful about

Two things.

First, the licence didn’t change. Qwen-Image-2.1’s weights are research-only, non-commercial. Quantising a model does not relicense it. Everything we said on the 20th still applies: excellent for personal work, learning and prototyping; a legal problem for anything you invoice. The quants inherit the terms of what they quantise.

Second, look at which repos are growing. Four of the five fastest-moving Qwen repos in our tracker have “Uncensored”, “Heretic” or “Abliterated” in the name. That is a real signal about what a chunk of the local-image-generation audience wants from open weights, and it is worth naming plainly rather than pretending the adoption curve is purely about VRAM headroom. The 8GB story is true. It is not the only story in the numbers.

Why this pattern keeps repeating

The gap between “weights released” and “weights usable” used to be weeks. For Qwen-Image-2.1 it was about 72 hours, and it was closed by unaffiliated people with quantisation toolchains, not by the lab that trained the model.

That’s the current shape of open-weights releases: the lab ships the artifact, and an informal supply chain of quantisers, node authors and workflow packagers decides within days whether it becomes something people actually run. Our tracker sees that supply chain more clearly than it sees the releases themselves — which is arguably the more useful thing to be watching.