XR / Spatial Computing

ARS-Avatar Learns Where a Body Shadows Itself, Which Is What Relighting Usually Gets Wrong

Differentiable screen-space ambient occlusion with per-body-part occlusion radii, optimised through finite differences — so an avatar can be dropped into new lighting and still look attached to itself.

Placing a photoreal person into a virtual environment convincingly is mostly a lighting problem. Geometry and texture are comparatively tractable; what breaks the illusion is that the captured human is still wearing the lighting of the room they were filmed in.

ARS-Avatar, posted 23 September 2026, tackles the specific reason that’s hard. The authors are Jiateng Liu, Hao Gao, Junxin Sun, Mengqi Liu, Jiu-Cheng Xie and Jucheng Song.

The entanglement problem

From the paper: creating animatable and relightable avatars from multi-view images is challenging because pose-dependent deformation, materials and light visibility are intrinsically coupled in images.

That’s the crux. A dark patch under an arm could be a shadow, a darker material, or geometry folding — and from pixels alone those are the same observation. Get the attribution wrong and the avatar’s shading stops responding correctly the moment you change the pose or the light.

What they do

Three pieces:

Surfel representation — surface elements rather than a mesh or volumetric field, for a high-quality animatable avatar built from multi-view images captured under unknown illumination. Not calibrated studio lighting; whatever light happened to be there.

Deformation priors from a template mesh, used as additional detail beyond the driving poses, to help estimate surfel attributes faithfully.

Deferred shading to estimate BRDF materials — separating what the surface is from how it was lit.

And the contribution that gives the paper its title:

A differentiable screen-space ambient occlusion formulation enabling gradient-based optimisation of body-part specific occlusion radii through finite differences — an efficient approximation of light visibility that can be jointly optimised with everything else.

Why learnable per-part occlusion radii is the clever bit

Ambient occlusion approximates how much ambient light reaches a point given nearby geometry. In real-time graphics it’s a screen-space hack with a hand-tuned radius, and one global radius is always wrong somewhere.

On a human body that’s acute. The occlusion scale appropriate to a torso — broad, gentle self-shadowing — is nothing like the scale for fingers, where geometry is small and self-occlusion is tight and frequent. A single radius gives you fingers that look flat or a torso that looks grubby.

Making the radius per body part and learnable lets the model discover the right scale for each region from data, instead of an artist guessing. And making it differentiable is what allows it to be optimised alongside materials and deformation rather than fixed beforehand — which matters, because those three quantities are the ones the paper says are entangled. You can only disentangle them by solving for them together.

Why this belongs under spatial computing

Because relighting is the requirement that separates an avatar you can show from an avatar you can place.

This week we covered Meta’s Hologram, which is not volumetric at all — a server-side generative video stream, driven by audio, that behaves as a billboard and rotates to face you. That approach works for a flat call and breaks when you walk toward it.

ARS-Avatar is the other end of the same problem: an actual geometric, relightable representation whose lighting responds to the environment you put it in. It’s research rather than a product, from multi-view capture rather than a phone, and it’s much closer to what “a person in your room” would need to be. The honest read is that consumer telepresence has taken the shortcut because the shortcut ships, while the work that would make it real is in papers like this one.