XR / Spatial Computing

Meta's Hologram Avatars Aren't Volumetric — They're a Diffusion Model Guessing From Your Voice

A hands-on found the Ray-Ban Display version convincing enough to mistake for a selfie feed, and the VR version breaking the moment you walk toward it.

We covered Meta VR Glasses yesterday and flagged Hologram — Meta’s photorealistic avatar system — as the headline software feature. A hands-on published September 24 fills in how it actually works, and the mechanism is not what the word “hologram” implies.

How you make one

You hold up a smartphone running the Meta AI app, perform some head movements, smile, and give long-form expressive answers to open-ended questions. The system captures your facial expressions, your clothing and your background. After processing, the Hologram is ready.

No capture rig, no studio, no volumetric stage. That’s a genuine achievement on the accessibility axis, and it’s the bar Apple set with Personas.

What it actually is

This is the part worth reading twice.

Holograms are not volumetric avatars. They are a real-time generative AI diffusion model running on Meta’s servers, taking in your microphone audio stream and producing a generated 2D video stream for Ray-Ban Display, or stereoscopic 3D video for the VR Glasses.

So:

  • The facial expressions are driven entirely by the caller’s voice. Not by face tracking. By audio.
  • The rendering happens on Meta’s servers, not on your device.
  • What the other person sees is synthesised video of a plausible you, inferred from how you sound.

That is a fundamentally different object from a tracked avatar. A tracked avatar is a puppet you are moving. A voice-driven diffusion Hologram is a model’s guess at what your face would be doing while saying what you’re saying. When you raise an eyebrow silently, there is nothing to transmit.

The hands-on findings

Honest results, and they split by device.

Ray-Ban Display (shoulders-up, for WhatsApp calls) was surprisingly realistic — one tester initially mistook a Hologram for a real selfie camera feed. That is a striking result, and it makes sense: a small, 2D, head-and-shoulders video is exactly the regime where generative video is strongest, and a phone-call framing sets low expectations for fidelity.

Meta VR Glasses (full-body) had clear problems:

  • Avatars appeared slightly undersized
  • Facial detail looked “quite blocky”
  • The illusion “completely broke when I leaned too far to the side or walked towards the Hologram, because it rotated to always face me”

That last one is the tell. A thing that always turns to face you is a billboard, not a body in your room. It’s the oldest trick in real-time graphics, and it’s what you use when you have a generated video plane rather than geometry. Walking around a person and having them swivel to track you is the single most presence-destroying thing a spatial avatar can do.

Timeline and the honest read

  • Ray-Ban Display: fall 2026, Early Access
  • Meta VR Glasses: spring 2027

The strategic logic is clear: ship the version that works now — flat, small, over WhatsApp — and let the hard version mature. Selling voice-driven generative video as telepresence on a phone call is defensible; nobody expects a WhatsApp call to be volumetric.

Selling it as a person standing in your room is a much bigger claim, and by the hands-on account it isn’t there yet. For anyone building spatial or performance work involving remote presence, the useful takeaway is that the leading consumer implementation of “realistic avatar” is currently a server-side video generator with no geometry, and it behaves like one at the edges. Plan for a billboard, not a body.