Last week we covered Breaking Taps designing a custom microprocessor and having it fabricated through wafer.space, and made the point that the ceiling on what an individual can build in silicon has moved.
Here’s the same shift from the other direction: a HackMIT-winning project that built a GPU in 24 hours, picked up by Adafruit.
Why this is remarkable and also not
The honest framing first, because “built a GPU” invites the wrong picture. This is not a competitor to anything with a fan on it. A hackathon GPU is a graphics pipeline implemented in hardware description language, running on an FPGA — geometry in, rasterised pixels out, at modest resolution and speed.
What makes it notable is the timescale. Implementing a rasteriser in hardware has historically been a serious undertaking: a semester of a graduate course, or a long personal project. The pipeline is genuinely involved — vertex transformation, clipping, triangle setup, edge functions, depth testing, framebuffer output — and each stage is a state machine you have to get right in a language that punishes software habits.
Doing it in a day means the surrounding tooling has improved enormously: modern HDLs, fast simulation, open synthesis toolchains, cheap FPGA boards with enough logic and memory, and — realistically — AI assistance for the boilerplate. That’s the story.
Why a creative-technology audience should care about GPUs specifically
Because the GPU is the one piece of the stack creative coders use constantly and understand least.
Every shader you write, every particle system, every Gaussian splat renderer runs on a machine whose behaviour you infer from performance rather than from knowledge. Why is a branch in a fragment shader expensive? Why does changing texture access patterns transform your frame rate? Why do draw calls cost more than triangles? Those questions have concrete answers in hardware, and the answers change how you write shaders.
Building even a toy rasteriser teaches the model. You discover for yourself that:
- Work is organised in lockstep groups, which is why divergent branches cost you
- Memory bandwidth, not arithmetic, is usually the limit
- Fixed-function stages exist because some operations are cheap in silicon and expensive in general-purpose logic
- Per-draw-call setup is real hardware state being reconfigured
None of that is mysterious once you’ve implemented it badly yourself.
The progression worth noticing
Over the last fortnight this beat has produced: CERN releasing a formally verified VHDL library, Breaking Taps taping out a custom microprocessor, the ESP32-S31 becoming capable of booting Linux, and now a GPU in a weekend.
That’s a coherent trend rather than four unrelated stories. The hardware design stack is becoming accessible at roughly the rate the software stack did twenty years ago — better languages, better free tools, cheaper targets, and reusable verified components. The practical consequence for creative work is that “design a custom digital circuit for this” is moving from a specialist commission to a thing an ambitious individual can attempt.
Not that most should. But the wall is somewhere different than it was.