Hand tracking is the input method everyone in XR has decided on — Apple built Vision Pro around gaze-and-pinch, Meta’s VR Glasses ship with no controllers at all, and the $1M developer competition Meta opened last week requires hands-only interaction.
A paper posted 25 September 2026 is a useful counterweight, because its starting premise is that common XR control methods, such as motion controllers and hand gestures, are still perceived as unintuitive. The authors are Alicia Torc, Carl Tornberg, Eric Piette, Renaud Ronsse and Benoit Macq.
The interface
The context is a real deployment problem: service robots are hard to put in domestic environments, partly because full autonomy isn’t reliable in unpredictable surroundings, and partly because conventional control methods are inaccessible to novice users. XR helps by letting an operator see robot information overlaid on the real world — but only if the interaction works.
Their interface uses a commercial XR pen to command a semi-autonomous mobile robot in AR:
- Point at a position in the room
- Select it
- Drag an augmented arrow to set the robot’s desired orientation at that destination
Two additional interfaces are evaluated for comparison.
Why a pen is a reasonable answer
The gesture-versus-stylus argument tends to be conducted on aesthetics — controllers feel like gaming, hands feel like magic. The practical case for a pen is more specific.
A pen has a tip, so pointing is unambiguous. Pointing with a finger requires the system to infer a ray from a hand pose, and small angular errors become large positional errors across a room. A stylus has a physical, visible, precisely tracked point.
It has a click, so selection is discrete. The hardest part of hand tracking is knowing when the user committed. Pinch detection is probabilistic, produces false positives, and gives no tactile confirmation. A button is a button.
It supports drag as a continuous, sustained gesture. Dragging an arrow to set orientation requires holding a selection while moving precisely — exactly the case where a pinch drifts or drops.
People already know how to hold one. The learning curve for a stylus is approximately zero, which matters enormously for the novice users this paper is explicitly about.
And the task suits it. “Specify a position and an orientation on a floor plan” is fundamentally a 2D drawing task performed in 3D space. Drawing is what pens are for.
The wider point for XR work
There’s a real tension worth naming between where the industry is going and what studies like this find. Hands-only is being chosen for product reasons — no accessory to lose, no cost, no charging, no barrier to first use — and those are legitimate. But they are not the same as being the best input method for precise spatial tasks.
For anyone building XR work: the useful question isn’t “hands or controllers” but what precision does my interaction actually need? Gaze and pinch is fine for selecting a large target. It is not fine for placing something exactly, drawing, or specifying an angle. If your piece needs precision, the honest options are to design the precision out of it, or to accept a physical input device.
The robotics framing shouldn’t put creative readers off. Specifying a position and orientation in a real room, in AR, using a tracked stylus is the same interaction as placing a virtual object in an installation, aiming a projection, marking up a space during install, or authoring a spatial piece in situ. The robot is just the thing that obeys.
Related Reading
- Comparative Evaluation of an XR Pen-based Control Interface for Semi-Autonomous Mobile Robot Navigation in Service Environments — arXiv:2609.31117
- Evaluating the Impact of Adaptive Extended Reality on Human-Robot Interaction — arXiv:2609.31138
- Through Human Eyes and Machine Eyes: view mismatch in video see-through XR — arXiv:2609.29173
- Meta VR Start Developer Competition — Devpost
- arXiv cs.HC — Human-Computer Interaction listings