Digital Artists

Replacing a Neural Network With a Pile of Blendshapes, and Getting 60fps on a Phone

GALA distils a 3D Gaussian avatar's expensive per-frame decoder into identity-independent blendshapes plus a shallow coefficient predictor. CPU animation cost drops by up to three orders of magnitude.

3D Gaussian avatars have an awkward asymmetry. Rendering them is fast — rasterising splats is exactly what GPUs are good at. Animating them is slow, because the shape of the avatar for a given expression comes out of a neural network that has to run every single frame.

So you have a representation that renders at hundreds of frames per second being fed by a decoder that cannot keep up with it. “One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars”, posted 1 October 2026 by Ramazan Fazylov, Stamatis Lefkimmiatis and Ivan Laptev, removes the decoder.

The method

GALA — Gaussian Animation via Linear Approximation. The insight: a pretrained avatar model can be approximated by linear combinations of identity-independent blendshapes.

Two components:

  1. The blendshapes are constructed via block-local PCA under rendering-aware constraints
  2. A shallow neural network predicts the blendshape coefficients

So instead of a large network producing thousands of Gaussian parameters per frame, you get a tiny network producing a handful of coefficients, which then index into a precomputed basis. The expensive work moved from runtime to a one-time distillation.

Why “blendshapes” is the right word and the right idea

Blendshapes are not new — they are how facial animation has worked in production for decades. You sculpt a set of target shapes (a smile, a brow raise, a jaw drop) and animate by weighting them. The FACS-derived rigs in every film and game pipeline are blendshape rigs, and Apple’s ARKit exposes 52 of them as its face-tracking API.

What makes this paper interesting is recovering that structure from a neural avatar rather than authoring it. The claim embedded in the method is that the neural decoder, despite being a nonlinear function with millions of parameters, is mostly doing something linear — and that the nonlinearity it adds is largely not visible once rendered.

That is a believable claim about faces specifically. Facial deformation is driven by a limited set of muscles acting roughly additively; the space of possible face shapes is genuinely low-dimensional. A big network trained on it will learn a function that a linear basis can approximate, because the underlying phenomenon is close to linear.

Two details in the construction are doing real work:

Block-local PCA, rather than global PCA. A single PCA over the entire avatar would mix unrelated regions into each component — a basis vector that moves the mouth and the back of the head together, because they happened to correlate in the training data. Doing it per-block keeps components spatially coherent, which both improves the approximation and makes the basis behave sensibly when you drive it with coefficients it never saw.

Rendering-aware constraints is the smarter half. Plain PCA minimises error in parameter space — it treats every Gaussian’s position, scale, rotation and colour as equally important numbers. But those parameters do not contribute equally to the image. A large, opaque, front-facing splat matters enormously; a tiny one buried inside the head matters not at all. Constraining the decomposition by its effect on the render rather than on the parameters spends the basis’s capacity where it is visible. This is the same insight that makes perceptual image codecs better than ones minimising mean squared error.

The results

  • CPU animation cost reduced by up to three orders of magnitude
  • 60fps on mobile devices
  • Most rendering quality preserved
  • Generalises to previously unseen identities without retraining the original models

Note that the headline reduction is in CPU cost. That is the correct framing and it tells you where the bottleneck actually was: the neural decode was running on the CPU (or competing for the GPU with the rendering), and it was the thing standing between a splat avatar and a phone.

Generalising to unseen identities is the result that makes this practical rather than a per-avatar optimisation. If every new person required redistilling a fresh basis, this would be a build-time step with a cost per character. Identity-independent blendshapes mean one basis serves everybody — hence the title.

What it means if you make things

Avatars on phones, standalone headsets and the web. Quest, Vision Pro and a browser are all platforms where a per-frame neural decode is the thing that kills a splat avatar. Replacing it with a basis lookup and a shallow net moves photoreal avatars into the budget those platforms actually have.

And the pattern transfers to anything with a neural decoder in a hot loop. The recipe is: your network is probably approximately linear over the manifold you care about; find the basis, fit a tiny predictor, move the cost to build time. That describes a lot of current real-time ML — neural materials, learned animation, deformation fields, audio synthesis models.

The honest caveat is in “most rendering quality preserved.” A linear approximation gives up the extremes. Expect the loss to show up on the least-linear parts of a face — the inside of a mouth, extreme asymmetric expressions, anything where skin folds and self-occludes. For a conversational avatar at normal expression ranges that is an excellent trade; for close-up dramatic performance it may not be.

Worth reading alongside the finding we covered yesterday that splat quality degrades more than metrics suggest when viewed stereoscopically. An efficiency win measured monoscopically deserves a look in a headset before you ship it.