Human activity recognition from wearables is a mature field with a clean benchmark story: new architecture, higher accuracy, publish. The benchmarks are stable, the metrics are agreed, and progress looks monotonic.
Characterizing the Performance Gap in Human Activity Recognition for Older Adults, posted 2 October 2026, asks who the progress is for. The authors are Hossein Khayami, Sungjin Hwang, Eshed Ohn-Bar, David E. Conroy, Amanda Lazar, Eun Kyoung Choe and Hernisa Kacorri.
The finding
Improvements on younger-adult benchmarks fail to transfer equally to data collected from older adults, resulting in a persistent and often widening performance gap.
Widening is the word that should stop you. It is not that older adults are served worse — that would be a static disparity, bad but stable. It is that the gap grows as the field advances, because every improvement is being selected for on a population that does not include them.
This is a well-understood failure mode with a precise mechanism: the benchmark is the objective. If your validation set is 80% people under 40, then architecture search, hyperparameter tuning and early stopping are all optimising for people under 40. An innovation that helps younger adults and hurts older ones scores as an improvement and gets adopted. Repeat for a decade and the gap compounds.
Why older adults’ movement is genuinely different
This is not a small distributional wrinkle. The signal is different in ways that matter for an IMU:
Gait changes substantially with age. Slower cadence, shorter stride, reduced hip and ankle range, less vertical displacement, longer double-support phase, more step-to-step variability. An accelerometer sees smaller, slower, noisier peaks.
Activities are performed differently. Sitting down becomes a controlled descent with hands on the chair arms rather than a drop. Stairs are taken one foot at a time with a handrail. Walking may involve a stick or a frame — which adds an entirely new periodic signal to the trace.
The activity set itself is different. Benchmarks are heavy on running, cycling and stair-climbing. The activities that matter for an older adult’s health monitoring are sit-to-stand transitions, slow ambulation, standing balance, and — critically — falls and near-falls, which are rare, safety-critical, and almost absent from datasets built by asking graduate students to wear a sensor.
And device placement is less consistent. A wrist-worn sensor on someone with reduced grip strength, or a waist clip on clothing that fits differently, produces different data from a lab-fitted rig.
So a model trained on young movement has learned an amplitude and frequency prior that older movement violates. It is not a harder problem; it is a different one, presented as the same one.
What worked, and the honest trade
Richer representations, particularly frozen self-supervised features pretrained on the age-diverse UK Biobank dataset, substantially improve performance for older adults and narrow the disparity, though at some cost to younger-adult accuracy.
Three things worth unpacking:
Self-supervised pretraining on a large, age-diverse corpus. The UK Biobank accelerometer dataset is roughly 100,000 participants across a wide age range, wearing wrist accelerometers for a week. Crucially it is unlabelled at the activity level — which is exactly what self-supervised learning wants. You pretrain on the raw signal, learning what human movement looks like in general, across bodies.
Frozen. The pretrained features are not fine-tuned. That is a deliberate choice and a revealing one: fine-tuning on a younger-skewed labelled set would drag the representation back toward the biased distribution. Freezing preserves the age-diverse structure and trains only a small head on top.
And the cost is stated. Younger-adult accuracy drops slightly. The authors say so rather than burying it, which is both honest and the correct framing — a model that is equitable across ages is not a free upgrade, it is a different operating point, and choosing it is a decision rather than an optimisation.
Their broader conclusion: genuine progress requires representations capturing population diversity and personalised adaptation to individual movement characteristics, not architectural improvements alone.
Why this matters for the work we cover
We have written about three IMU-based systems in the last fortnight — whole-body pose from a single earbud IMU, the SPAR boxing wearable, and SkeletonDance’s rhythm feedback for beginner dancers — and published an IMU primer. Every one of them rests on a model or heuristic tuned to some population’s movement.
The practical reading for anyone building movement-sensing work:
Your thresholds encode a body. A step-detection threshold, a gesture-onset trigger, a “has the person moved enough” gate — all of these were tuned on whoever was in the room when you tested. If that was you and three colleagues, your installation works for people your age and build.
Test with the people who will use it. For an interactive piece in a public venue, that includes children, older visitors, people with mobility aids and people who move cautiously because they are unsure what the thing does. The last group is a large fraction of any gallery audience and they all move differently from a confident tester.
And an adaptive baseline beats a fixed threshold. Calibrating to the individual in the first few seconds — their resting signal, their movement amplitude — is cheap and removes most of this class of problem without any machine learning at all.