Research & Innovation

Whole-Body Pose From the IMU in One AirPod

79mm lower-body error and 0.809 F1 on foot contact, from a single head-mounted sensor — and adding foot sensors made it worse.

Sparse-sensor motion capture has been converging on the same answer for years: fewer sensors than you think, in better places than you expect. “One Sensor, Whole Body — 3D Body Pose from a Single Consumer Earbud IMU”, posted 28 September 2026, pushes that to its limit and asks whether one is enough.

The authors are Zhilin Guo, Boqiao Zhang, Oszkár Urbán, Josef Bengtson, Hakan Aktas, Wenzhao Li, Siyu Hong, Kyle Fogarty, Chenliang Zhou, Ali Senguel and Cengiz Oztireli.

What they did

Built a capture system combining RGB-D video with IMU data from AirPods and foot sensors, producing a benchmark of 35 sequences across a range of motions. Then asked what a single head IMU can recover.

The numbers

  • 79.0 mm rigid-MPJPE on lower-body pose
  • 0.809 macro-F1 on per-foot ground contact

Rigid-MPJPE is Mean Per Joint Position Error after rigid alignment — the average distance between predicted and true joint positions once you have removed global translation and rotation. It measures body configuration, not where the body is in the world. 79 mm is roughly the width of a fist: not motion-capture-suit accurate, comfortably good enough to know how someone is standing, walking, crouching or turning.

Per-foot ground contact at 0.809 macro-F1 is the more practically loaded number. Knowing which foot is on the floor, and when, is what lets you place a figure in a world without it sliding — foot contact is the constraint every retargeting and physics-based animation pipeline is built around. Getting it from an earbud is a genuinely surprising result.

Why a head sensor can see the legs at all

Because walking is a whole-body oscillation and the head is at the end of the lever.

Every step sends an impulse up the kinetic chain. The head bobs vertically at step frequency, rolls slightly toward the stance leg, and pitches with the gait cycle. Heel strike is a sharp, unmistakable transient. The head does not move independently of the legs during locomotion — it moves as a consequence of them, and a model trained on that relationship can invert it.

The limits follow from the same logic. The paper is explicit: while the earbud captures some arm movement, it cannot effectively capture upper-body kinematics. Of course not — you can wave an arm without your head moving at all. Arms are decoupled from the head in a way legs are not. The result is a lower-body system, honestly labelled.

The finding worth remembering

Adding foot sensors did not improve accuracy — and sometimes degraded it, due to sensor quality rather than information loss.

That is a genuinely useful negative result, and it inverts the intuition that governs most sensor-fusion design. Foot IMUs are closer to the thing being measured; more information should not hurt. But a noisy, drifting, poorly-calibrated sensor fed into a fusion model does not add a little value — it adds a confident wrong signal that the model has to learn to distrust, and it does not always learn to distrust it enough.

The authors’ framing: reliability of individual sensors matters more than quantity. Anyone who has built a multi-sensor rig has learned this the expensive way, usually after a performance where one flaky sensor poisoned an otherwise working system.

Note that a companion paper from the same cluster — “Reliability-Gated Fusion of Consumer Head and Foot IMUs for Lower-Body 3D Pose”, posted the same day — appears to be the direct follow-up: if bad sensors hurt, gate on estimated reliability and only fuse what you trust.

Why this matters outside the lab

The sensor is already in people’s ears. That is the whole argument. Every previous sparse-IMU system required someone to put on hardware they would not otherwise wear. Earbuds are worn by an enormous number of people, continuously, for reasons that have nothing to do with motion capture. A pose estimate that needs no additional hardware is a different category of thing from one that needs a strap on each limb.

Concretely, for creative and interactive work:

  • Audience-scale interaction. A piece that responds to how people are moving through a space, without handing out sensors, is a fundamentally different proposition — and gait and contact are exactly what you would want for crowd-responsive sound or light.
  • Performance capture without a suit. Rough lower-body pose and reliable foot contact is enough to drive a stylised avatar or a generative visual, particularly for dance and walking-based work.
  • Spatial audio that knows your gait. Head-tracked audio already uses this IMU for orientation. Knowing the step cycle as well opens up rendering that is locked to footfall.

The caution, as always: 35 sequences is a benchmark, not a deployment. And the privacy shape of this is worth stating — the same sensor stream that reveals gait also identifies people by it, and that data path already exists on hardware most people never think of as a sensor.