XR / Spatial Computing

Looking at an AR Assistant Beats a Wake Word, but Nobody Notices When It Stops Listening

Gaze acquisition completed before speech in 490 of 496 attempts. The failure the study found is at the other end: release is invisible once you look back at your work.

Wake words carry a trade-off nobody has solved: raise the sensitivity and you catch more commands and more accidents; lower it and the reverse. In co-located augmented reality — where you are wearing a headset in a room with another person — there is a second problem on top. The system has to work out whether you are talking to it or to them.

Beyond Activation: Gaze Invocation with Visible Status for an Embodied AR Assistant in Co-Located Collaboration, posted 7 October 2026 by Chenrui Ma, Yoshio Ishiguro and Qing Zhang, replaces the wake word with where you are looking — and finds that this solves the easy half of the problem and exposes the hard half.

The design

Four ideas combined, and the combination is the contribution:

  • Gaze-based address. You look at the assistant to talk to it.
  • An embodied assistant anchored in the scene. Not an overlay or a voice from nowhere — an object with a location in the room.
  • Visible listening status on its body. Activation progress and listening state are rendered on the thing itself.
  • Turn management across all three phases — activation, continued interaction, and release.

Mechanically: sustained gaze opens a local interaction channel; speech and playback keep it open when your attention returns to the task; inactivity closes it.

The gaze-as-address idea is well founded. Looking at someone before speaking to them is how humans handle addressing in a group, it requires no learned command, and critically it is a signal the other person in the room can also read — your collaborator can see you turn to the assistant, which is exactly the disambiguation a wake word fails to provide.

Entry works

25 participants. Each chose the assistant’s placement and dwell duration, then worked with a partner to plan a trip while using it. The study recorded 500 interaction outcomes, including 496 assistant-directed utterance attempts.

Gaze acquisition completed before speech in 490 of those 496 attempts — 98.8%. Of those 490, 467 began while the channel was already open; 23 began after it had been released. Only 6 attempts started before acquisition finished.

That is a strong result. People looked at the thing before talking to it, naturally, without being drilled — which is the whole bet, and it paid.

Release is where it falls apart

Here is the finding that makes the paper worth reading rather than just citing.

Of 102 reviewed attempts in which gaze left after acquisition but before speech, the retention policy kept the channel open for 79 and released it before 23.

And the authors’ own conclusion: an embodied target with visible status supports clear entry, while revealing a different problem at release — after users look back at their work, they may not see that the assistant has stopped listening.

This is a structurally awkward problem, and it is worth being precise about why.

The status display lives on the assistant. That is the right place for it during entry — you are looking at the assistant, so you see the activation progress, and the feedback is exactly where your attention already is.

But release happens by definition when you are not looking at it. The channel closes on inactivity, which means it closes while your eyes are on the task. The one piece of state you most need to know — am I still connected — is rendered in the one place you are guaranteed not to be looking.

So you speak into a closed channel. Nothing happens. And unlike a wake-word system, there is no clear action to recover with, because there was no explicit command to repeat — you have to notice the failure, look back, re-acquire, and start again.

Why this matters beyond assistants

Anyone building gaze-driven interaction in XR will hit the same shape of problem, and it is not specific to voice.

Gaze as input has an inherent asymmetry: it is also your primary perceptual channel. A mouse cursor is not where you look; a controller is not where you look. Gaze is both the pointer and the eye. Any system that consumes gaze as input is competing with the user’s need to look at something else, and any feedback that requires looking at a specific place is unavailable precisely when the user has moved on.

The implication for designers is uncomfortable but clear: state that must survive a shift in attention cannot be displayed at the point of attention. It has to be peripheral, or ambient, or audible, or haptic — in a modality that does not require the user to aim at it. The authors frame their contribution partly as exactly this: design implications for communicating assistant state after visual attention moves elsewhere.

The paper is 11 pages, 9 figures, 2 tables, classified H.5.1 and H.5.2.