Research & Innovation

A New Research Framework Wants Your Coffee Mug to Be Part of the Gesture

Objestures merges everyday-object interaction with mid-air gesture control into one design framework, tested across 30 real application examples — an attempt to finally unify two interaction paradigms researchers usually study apart.

A lot of interaction-design research treats “everyday objects as interfaces” and “mid-air gestures” as two separate research tracks, each with its own narrow set of demos and its own specific hardware assumptions. A recent paper, Objestures: Everyday Objects Meet Mid-Air Gestures for Expressive Interaction, argues that treating them separately has been leaving a real design space unexplored — the space where you pick up a mug and gesture with it, rather than choosing one interaction mode or the other.

What the framework actually proposes

Objestures introduces five interaction types that combine object-based and mid-air gesture approaches into one design vocabulary, rather than a single new gesture set. The researchers tested it two ways: an exploratory study with 12 participants performing 3D rotation and scaling tasks, and case studies across three built applications — Sound, Draw, and Shadow — evaluating how the combined interactions actually felt to use.

What the results actually showed

On basic 3D manipulation tasks, performance using Objestures-style interactions matched what a headset’s native freehand gesture system already achieves on its own — meaning the object-plus-gesture combination isn’t a regression from pure freehand manipulation, it’s an addition on top of it. Participants described the resulting interactions as intuitive, engaging, and expressive, with real interest in using something like it in everyday contexts rather than only in a lab demo. The researchers also demonstrated the framework’s range across 30 real-world application examples, arguing for its versatility rather than resting on the three built case studies alone.

Why this is worth translating for a working audience

Most gesture-interface products ship a full freehand system, or ship an object-tracking system, rarely both combined with any shared design logic. A framework that gives designers concrete guidance for building uni- and bimanual interactions that blend the two — rather than treating “hold an object” and “gesture in the air” as competing choices — is directly useful groundwork for anyone building spatial-computing or installation interfaces that want a hand to carry meaning whether or not it’s holding something.