Research & Innovation

A Diffusion Model Running on a Microcontroller With No Linux and No GPU

64×64 grayscale images from a 3.36MB quantised model on an STM32N6570-DK, and a 289M-parameter language model on an NXP FRDM-MCXN947. Both open source.

The interesting question about generative models has stopped being how large they can get. Everyone knows the answer is “larger.” The interesting question is how small they can get and still do something, because that number determines what can be built into an object that runs by itself.

Two open-source projects, reported by Ananthu Ashok for Open Source For You, have pushed that floor down to bare-metal microcontrollers — no Linux, no GPU, no external accelerator.

Project one: a diffusion model on an STM32N6570-DK

  • Board pairs an Arm Cortex-M55 with an Ethos-U55 neural processing unit
  • Generates 64×64 grayscale images
  • Quantised model occupies about 3.36 MB of memory

The hardware choice is the story. The Cortex-M55 is the first Armv8.1-M core with Helium (M-Profile Vector Extension) — SIMD in a microcontroller class part. The Ethos-U55 is a microNPU, a small fixed-function accelerator for quantised integer neural network operations, designed to sit next to an M-class core rather than replace it.

That pairing is Arm’s answer to edge inference, and it has been shipping for a while. What is new is somebody running a diffusion model through it, which is a far heavier workload than the keyword-spotting and image-classification tasks these parts were pitched for. Diffusion is iterative — you run the network once per denoising step, so twenty steps means twenty full forward passes. On a microcontroller, even at 64×64, that is a meaningful amount of arithmetic.

3.36 MB is the number to hold onto. That is the entire model, quantised, in flash. For comparison, Stable Diffusion’s UNet at fp16 is well over a gigabyte.

Project two: 289M parameters on an NXP FRDM-MCXN947

A 289-million-parameter language model on an NXP microcontroller, using heavy quantisation, compact weight storage and optimised inference to fit the available memory and compute budget.

289M is small by current standards and not small in the history of the field — it is roughly GPT-2 Medium’s size, and GPT-2 was a 2019 research result that people found startling. Running that class of model on a part costing a few tens of dollars, with no operating system underneath it, is the kind of thing that would have read as implausible five years ago.

The honest framing, which both projects supply

Neither project claims microcontrollers can stand in for GPU-based systems. That disclaimer is in the reporting and it is correct. A 64×64 grayscale image is a proof of capability, not a product. A 289M model is not going to hold a conversation you would enjoy.

What they demonstrate is different and more useful: the compression and hardware-aware optimisation work needed to make models fit. And both are open source, so that work is available to build on — which is the actual deliverable. The image is a demo; the quantisation and weight-packing pipeline is the contribution.

Why this matters for making things

Because “no Linux” is a completely different engineering object from “small Linux.”

A Raspberry Pi running a small model is a computer. It boots in tens of seconds, needs a filesystem that can be corrupted by yanking power, wants a few watts, and has an operating system that will eventually want updating. For an installation that has to survive a gallery attendant flipping a breaker every night for four months, all of that is risk.

A bare-metal microcontroller starts in milliseconds, cannot corrupt a filesystem it does not have, runs on a fraction of the power, and has no software that can rot. Power it from a battery or a solar cell and it is a genuinely autonomous object.

So what becomes possible:

  • Self-contained generative objects — a small screen, a battery, and a model that makes images with no network and no host computer. The work is the object, not a client for a service elsewhere.
  • Generative work that runs for years. No API to be deprecated, no subscription, no vendor. The piece in ten years does exactly what it does now, which is a claim almost nothing else in AI art can make.
  • Instruments with local intelligence — a synth, a controller, a wearable that runs a model in its own firmware rather than phoning home.
  • Editions. A limited run of physical objects each containing its own generative model is an entirely different proposition from a limited run of prints, and the economics work at these part costs.

The constraint is severe and the constraint is the point. 64×64 grayscale is a legitimate aesthetic — it is roughly a Game Boy screen — and work made at that resolution by a chip that can be embedded in anything is a different kind of object from a 4K render made in a data centre.