Research & Innovation

Seeing Speech Animates Faces by Modelling Three Motions: Spreading, Opening and Protrusion

Speech-driven facial animation keeps improving vertex accuracy while still looking wrong. This decomposes visible articulation instead.

Speech-driven facial animation has a characteristic failure that metrics don’t catch. The vertices are in roughly the right places, the timing is right, and the mouth still looks like it’s miming rather than speaking.

Seeing Speech, posted 24 September 2026, diagnoses why and proposes a structural fix. The authors are Hyung Kyu Kim, Byungchan Hwang and Hak Gu Kim.

The two problems

The paper names both precisely:

Speech production follows structured and constrained articulator coordination. Lips, jaw and tongue don’t move independently — they move in coordinated patterns dictated by what’s physically required to produce a sound. A model that treats the face as a free-form mesh isn’t wrong so much as unconstrained, and it will happily produce combinations of movements no human mouth makes.

The mapping from acoustics to motion is inherently one-to-many. The same sound can be produced by different articulator configurations, and the same configuration can produce different sounds depending on context. A model trained to regress motion from audio is being asked to pick one answer to a question with several correct ones — and regression under ambiguity averages, which is exactly why generated speech animation looks under-articulated and mushy.

The decomposition

Rather than predicting vertices directly, the framework models visible speech through three directional articulatory motions:

  • Spreading — lateral, as in “ee”
  • Opening — vertical, as in “ah”
  • Protrusion — forward, as in “oo”

These are composed into surface-consistent 3D facial motion.

To connect sound to those motions they propose a Speech–Articulatory Memory (SAM), capturing the correspondence between speech and the three motions under phonetic context.

Why three axes is the right abstraction

This is the bit worth appreciating, because it’s borrowed from linguistics rather than invented.

Phonetics has long described vowels along roughly these dimensions — height (open/close), backness, and rounding — and consonants by place and manner of articulation. Spreading, opening and protrusion is very close to a graphics-friendly restatement of the vowel space.

That matters for two reasons. It’s a low-dimensional, physically meaningful basis, so the model can only produce combinations that correspond to real articulation — which addresses the unconstrained-mesh problem directly. And phonetic context is exactly what disambiguates the one-to-many mapping: the visible shape of a sound depends on the sounds around it, which is coarticulation, and a memory indexed by phonetic context is a reasonable way to encode it.

Where it’s useful

Avatars and virtual performers. This is the visible half of the problem Meta’s Hologram takes a shortcut on — generating video from audio rather than driving geometry. A model that drives actual facial geometry from speech, with articulation that survives scrutiny, is what a geometric avatar needs.

Animation production. Automatic lip sync exists and animators routinely fix it by hand, because the automatic version is under-articulated in the way this paper describes. Better articulation means less correction.

Anything multilingual. A phonetically grounded representation should transfer across languages better than one trained to map audio to vertices for one language’s data, since the articulatory basis is universal even when the phoneme inventory isn’t.

It also sits neatly alongside the week’s other speech work — the backchannel timing paper and Audio8’s semantic VAD. Three different groups, all concluding that the useful move in speech interfaces is to model the structure of speaking rather than treating audio as an undifferentiated signal.