Speech-driven facial animation has reached the point where it is clearly working and clearly not right. The lips hit the phonemes, the timing is correct, nothing is obviously broken — and it reads as slightly lifeless in a way that is hard to name and therefore hard to fix.
The Shape of Speech: A Geometric Measure of Coarticulation for Speech-Driven 3D Facial Animation, posted 2 October 2026 by Danzel Serrano and Przemyslaw Musialski, names it.
The measure
Lip-path length compared with the shortest route through the vowel, consonant and vowel positions of a speech segment.
Take a VCV segment — vowel, consonant, vowel, the basic unit of coarticulation study, like the a-b-a in “aba”. Each of those three sounds has a canonical lip configuration. Draw the straight-line path through those three positions: that is the shortest route, the minimum the lips must travel to hit all three targets.
Then measure how far the lips actually travel. The ratio is the measure.
Real speech travels further than the direct path. The excess is coarticulation.
What coarticulation is and why the excess exists
Coarticulation is the influence sounds exert on their neighbours. You do not produce phonemes as discrete states; your articulators are continuously in motion, and the configuration for any given sound is modified by what came before and what is coming next. The /k/ in “key” and the /k/ in “coo” are made at different places in the mouth because of the following vowel — your tongue is already heading there.
So the lips overshoot, anticipate and curve rather than moving in straight lines between targets. The path is longer than it needs to be because articulation is a physical system with mass and momentum pursuing moving goals.
This is also why real speech looks alive. The excess travel is the fast articulatory component — the quick, high-curvature movement between sustained positions.
The finding
All four systems tested trace flatter lip trajectories than captured speech, with deficits equivalent to removing 15–60% of real speech’s fast articulatory component.
Fifteen to sixty per cent. Not a marginal shortfall — in the worst case, more than half of the fast movement is simply absent.
And it was validated perceptually: viewers strongly prefer real speech.
Why every system under-moves, structurally
This is the useful part, because the mechanism explains why it is not one system’s bug.
Regression to the mean. Almost all of these models are trained to minimise error against ground truth — L2 on vertex positions or blendshape weights. A model uncertain about exactly how far the lips will overshoot minimises expected squared error by predicting the average, which is less movement than any individual instance. L2 loss systematically produces under-articulation, for the same reason it produces blurry images in generative models.
Temporal smoothing. Pipelines smooth output to avoid jitter. Smoothing is a low-pass filter, and the fast articulatory component is by definition the high-frequency part. You are explicitly removing the thing that is missing.
Phoneme-target thinking. Systems built around hitting canonical visemes are modelling speech as a sequence of states with interpolation between them — which is the straight-line path. Coarticulation is precisely the deviation from that model.
And nobody was measuring it. Standard metrics are vertex error and lip-sync accuracy. Both reward hitting targets; neither notices that you took the short way round. A metric that is blind to a deficiency will not prevent it.
What to do with it
It is a trainable objective. The strength of this paper is that it hands you a target. A path-length ratio is differentiable and cheap, which means you can add it as a regularisation term: penalise a model whose predicted trajectories are flatter than the data’s. That is a much more direct fix than more training data.
Stop smoothing so hard. If your animation pipeline has a smoothing pass, measure what it is costing you. Jitter and articulatory speed live in the same frequency band, and a filter cannot tell them apart.
For anyone using these tools — and a lot of this audience is, for virtual performers, game dialogue, avatar work — the practical read is that the problem is under-articulation, not bad timing. If your character looks dead, the instinct is to blame the lip-sync. The actual fix is more movement: exaggerate the blendshape range, reduce smoothing, or drive it harder than feels correct. Animators have known this empirically for decades, which is why hand-keyed dialogue is pushed well past life; this paper explains why.
And it is another entry in a pattern we keep running into. Text-to-3D evaluation where the protocol varies more than the generators. Splat quality that monoscopic metrics do not penalise. Gesture accuracy that does not track naturalness. Four papers in a fortnight, all finding that the metric the field optimises is not measuring the thing that matters. That is starting to look less like a coincidence and more like the current state of graphics and HCI evaluation.