Emerging Interfaces

Generated Pointing Can Be More Accurate Than Human Pointing and Still Look Wrong

A benchmark of 2,000 annotated clips from VR conversations, with 3D ground-truth object positions — and a finding that geometric accuracy and perceived naturalness come apart.

Avatar gesture generation has got good at producing plausible motion — hands that move with the rhythm of speech, in the register of a person talking. What it has not had to do is mean anything spatially.

A Benchmark for Spatially Grounded Gesture Generation, posted 2 October 2026 by Anna Deichler, Rishabh Dabral, Fethiye Irmak Dogan, Anindita Ghosh and Jonas Beskow, is about the gestures that do.

Why pointing is the hard case

The authors’ framing: “pointing gestures anchor language to the environment.”

That is a precise claim about what deixis does. When you say “put it over there”, the sentence is incomplete — the referent lives in your hand, not in your words. A pointing gesture is the only part of the utterance that carries the location, which means a wrong point is a wrong sentence, in a way that a badly-timed conversational beat gesture is not.

So pointing is the case where gesture generation stops being a question of plausibility and becomes a question of correctness. Either the hand indicates the thing or it does not.

The benchmark

  • ~2,000 annotated clips from naturalistic VR conversations
  • 3D ground-truth object locations
  • Three separate dimensions measured:
    1. Temporal alignment — timing of gestures within speech
    2. Spatial grounding — accuracy of pointing to the intended referent
    3. Perceived naturalness — how human-like the gestures appear
  • A baseline model, MM-Conv-Flow, compared against human motion and a retrieval-based system

Using VR conversations as the capture method is the smart part. The problem with studying pointing is that you need to know, exactly, where the thing being pointed at is — and in a real room that requires tracking every object. In VR the scene is synthetic, so 3D ground truth is free and exact. You get naturalistic conversation with perfect annotation of the physical referent, which is not obtainable any other way.

The finding

Geometric grounding can exceed that of human pointing without any gain in perceived naturalness.

A system can point more accurately than a person does and look no more natural for it.

Which means the two things are not on the same axis, and that has a specific implication: you cannot optimise one metric and expect the other to follow. The authors’ conclusion — that referential gesture quality cannot be assessed through a single metric — is the paper’s actual contribution, more than the baseline model.

Why accuracy and naturalness come apart

Worth thinking through, because it tells you what naturalness is made of.

Human pointing is sloppy, and the sloppiness is informative. People point approximately and let language and context do the rest — “the red one, over there”. Precision is expensive and usually unnecessary. A gesture that is too accurate reads as mechanical, in the same way that perfectly quantised drumming reads as a machine.

Human pointing is a whole-body event with a specific temporal shape. There is anticipation, a stroke, a hold at the extreme, and a retraction — and the hold is what makes it readable. A system optimising for angular accuracy will drive the hand to the correct bearing and may skip the preparation and the hold entirely, which is geometrically perfect and gesturally illegible.

And gaze goes with it. People look at what they point at, usually slightly before the hand arrives. An arm that is correct while the head faces forward is uncanny in a way that no pointing-accuracy metric can see.

What this means if you build with avatars

Measure both, and expect to trade. If you are generating gesture for an avatar — in VR, in a telepresence system, for a virtual performer — you need an accuracy measure and a perceptual one, and the benchmark exists now so you can.

Deixis is where avatar systems break most visibly, and it is under-tested. Conversational gesture generation is evaluated on naturalness almost exclusively, because until you put the avatar in a shared space with objects there is nothing to be accurate about. The moment you do — collaborative VR, a guided tour, a virtual instructor pointing at a thing — the requirement changes category.

And the honest reading of the result is that naturalness is a composite of timing, hold, preparation, gaze and body involvement, most of which are not pointing accuracy. A system that nails the bearing has solved the easy half.

This connects to something we wrote about last week: SkeletonDance found subjective benefit without consistent objective improvement, and here the asymmetry runs the other way — objective improvement without subjective gain. Both are arguments for the same methodological point, which is that movement systems need to be measured on at least two axes because the axes genuinely diverge.