Kinetic typography has an awkward technical problem underneath its very pleasant results. If you let software deform a letterform freely — by optimisation, or by a generative model predicting vertex positions — you get tearing, self-intersection and jitter. The outline crosses itself, the counters collapse, and the letter stops being legible. If you constrain it tightly enough to avoid that, you get motion so stiff it may as well be a transform.
“Strike a Chord! Modal Kinetic Typography”, posted 29 September 2026 by Maham Tanveer, Jiyeon Han, Nanxuan Zhao and Hao Zhang, solves it by asking what the letter would do if you hit it.
The method
The goal: animate a vector glyph to express a semantic concept while keeping it legible.
The approach:
- Finite-element analysis extracts the letter’s softest vibrational modes — both for the whole letter and for its individual components
- A frozen video diffusion model controls only the amplitudes and phases of those modes
That is it, and the elegance is in the second step. The diffusion model is not allowed to move points. It is allowed to decide how much of mode 1 and mode 2 and mode 7, and when. Which restricts motion to smooth, seamlessly looping patterns and makes outline tearing structurally impossible.
What “modes” means and why physics is the right source
Hit a bell and it rings at particular frequencies. Those are its normal modes — the natural ways its geometry can deform, each with a characteristic shape and frequency. Every shape has them, determined entirely by its form and material. Modal analysis is the standard engineering technique for finding them, and finite-element analysis is how you compute them for an arbitrary geometry.
The softest modes are the lowest-frequency ones, and those are the ones that produce large, smooth, global deformation rather than high-frequency local wobble. A letter’s softest modes are things like: the stem bending, the bowl breathing, the crossbar flexing. Exactly the motions a designer would hand-animate.
Two properties make this a better basis than anything learned:
It comes from the glyph, not from a category. The authors are explicit that the modes emerge from the geometry itself rather than relying on category-specific priors like skeletons or keypoints. A skeleton-based animation system needs to know what kind of thing it is animating. An ‘S’ has no skeleton anyone agreed on. But it has vibration modes, and so does a ‘W’, and so does a logotype, and so does a Devanagari character. The method does not need to be told what a letter is.
Linear combinations stay valid. Mix two modes in any proportion and you get a physically plausible deformation, because modes are the eigenvectors of the system — they are orthogonal and they superpose. This is why tearing is impossible: you are not moving points, you are weighting shapes that were already smooth.
The pattern this belongs to
This is the second paper in two days built on the same idea, and the coincidence is worth naming because the idea is clearly having a moment.
Yesterday we covered GALA, which distils a neural Gaussian avatar decoder into blendshapes — a basis built by PCA — and has a shallow network predict only coefficients. Today’s paper builds a basis by FEA modal analysis and has a diffusion model predict only amplitudes and phases.
Same structure: find a low-dimensional basis of valid configurations, then let the model choose a point in that space rather than an arbitrary output.
The benefits are the same in both cases and they are substantial:
- Invalid outputs become unrepresentable. You cannot tear an outline if you cannot move a point. You cannot produce an impossible face if every face is a weighting of plausible ones.
- The model gets much smaller. Predicting a dozen coefficients is a far easier job than predicting thousands of vertex positions, which is why a frozen diffusion model suffices here with no retraining.
- The output is controllable and editable. Coefficients are a handle a human can grab. Vertex deltas are not.
For anyone building generative tools, that is the transferable lesson: constrain the output space with domain structure instead of asking a larger model to learn the constraints. It is cheaper, it is more reliable, and it leaves the artist something to adjust.
What you would do with it
The obvious applications are title sequences, logo animation, motion graphics and anything where type needs to behave like its meaning — a word for “wobble” that wobbles, “melt” that melts, while remaining readable.
The less obvious one: this is a general vector-animation method that happens to be demonstrated on type. Any closed vector outline has modes. An icon set, a map, an illustration, a UI element, a logomark — all animatable this way, all guaranteed not to self-intersect, all looping cleanly.
There is a project page with results, which is the right place to judge whether the motion actually reads. This kind of claim is one you have to watch rather than read.
We have covered the broader state of generative kinetic typography before; this is the most structurally interesting approach to it we have seen.