XR / Spatial Computing

AR Captions That Tell You Who Said It, From People Introducing Themselves

VocalEyes builds speaker voice profiles out of the self-introductions that already happen at the start of a meeting — no enrolment step, 88% identification, and speaker-tracking up from 47% to 87%.

Live captioning has got very good at the hard acoustic problem and remains bad at a social one. You can read every word in a four-person conversation and have no idea who said which.

For anyone relying on captions, that is not a cosmetic gap. Attribution is how you know whether a statement was a proposal or an objection, whether someone is repeating themselves, whether the room agrees. A transcript without speakers is a much weaker artefact than one with them.

VocalEyes: Speaker-Aware Augmented Reality Captioning through In-Conversation Registration, posted 26 September 2026 by Yuxiao Wang, Xulong Tang, Chen Chen and Rawan Alghofaili, attacks the attribution problem specifically.

The idea: register speakers from the conversation itself

Speaker identification normally requires enrolment — each person reads a passage beforehand so the system can build a voiceprint. That works in a lab and essentially never happens in life. You do not ask five people to record calibration audio before a meeting starts.

The alternative is anonymous clustering: group the audio into “Speaker 1, Speaker 2, Speaker 3” without knowing who they are. Standard diarization does this well, and it produces captions attributed to numbers, which is only marginally better than no attribution at all.

VocalEyes’ move is to notice that the enrolment data already exists in the conversation. Meetings begin with people introducing themselves. That introduction is a labelled audio sample: a voice, plus that voice stating its own name.

So the system builds identification profiles from natural self-introductions during the meeting. No enrolment step, and no anonymous numbers — actual names, obtained the way humans obtain them.

The interface, and the results

The full interface combines speaker-attributed captions, profile cards, and visual cues marking the speaking person’s face.

MeasureResult
Speaker identification accuracy88.0%
Speaker-tracking accuracy, caption-only AR47.2%
Speaker-tracking accuracy, full interface87.3%
Cognitive workloadLower with the full system

47.2% → 87.3% is the number that matters, and it is worth reading carefully because it is not the same as the 88% identification figure.

Identification accuracy is whether the system got the speaker right. Speaker-tracking accuracy is whether the user knew who was talking. Caption-only AR — words floating in your field of view, correctly transcribed — left people right about half the time, which is barely better than guessing between a handful of participants.

In other words: the model already knowing who spoke does not help unless the interface communicates it. Marking the speaking person’s face is what closed the gap. That is a design finding, not a machine learning one, and it is the more transferable of the two.

Why “lower cognitive workload” is the most important row

Because it is the thing that decides whether anyone uses this for more than a demo.

Reading captions while tracking a conversation is effortful. You are visually parsing text, mapping it onto people, maintaining conversational state, and formulating your own responses, all at once. AR adds to that load before it subtracts — there is now content in your visual field competing with the faces you are trying to read.

A system that improved accuracy while increasing workload would be a worse tool. Measuring workload alongside accuracy is the correct evaluation and it is skipped surprisingly often.

What to take from this beyond accessibility

Registration from natural behaviour, rather than a setup step, is a general design pattern and it is underused. The question to ask about any calibration requirement is: is the information I need already present in something the user does anyway?

  • Gaze calibration that happens while someone reads a welcome screen
  • Room scanning that happens while someone walks to their seat
  • Hand-size calibration from the first natural grab
  • Instrument or controller calibration from a warm-up rather than a test tone

Every setup step is a place people drop out, and the ones that feel like homework are the worst offenders. An interaction designed so that setup is invisible has a fundamentally different adoption curve.

And the attribution problem generalises beyond captions. Any system that renders multiple sources into one channel loses source identity — multi-channel audio mixed to stereo, several sensors into one visualisation, several collaborators into one document. The fix is usually the same as VocalEyes’: keep the identity and spatialise it back onto the source, rather than labelling it in text.

The obvious caution: this is a small-group result — the paper scopes itself to small-group discussions — and self-introductions are a Western meeting convention, not a universal one. But the mechanism is sound and the interface finding stands on its own.