The fastest way to make a piece of AR look fake is to leave occlusion out. A virtual object placed behind a real sofa draws on top of the sofa, and the brain rejects the whole scene instantly — not because the lighting is wrong or the shadows are missing, but because occlusion is the depth cue the visual system trusts most and the one it cannot be talked out of.
The WebXR Depth Sensing module fixes this, and the core of the fix is about four lines of shader.
Grae n walking through depth occlusion in a browser-based AR scene.
What the API gives you
Each frame, a depth map of the real world from the viewer’s perspective — a low-resolution image where each value is the distance to the nearest real surface along that ray. On a headset with depth hardware this comes from the device’s own reconstruction; on phones it may be derived from a time-of-flight sensor or from monocular estimation.
Two critical properties up front:
The resolution is much lower than your render target. Typically on the order of a couple of hundred pixels on the long edge. You are not getting a per-pixel depth image.
Support is uneven. It is available on Quest browsers and on Android Chrome with ARCore depth, and not on iOS Safari. Feature-detect and degrade; never assume it.
Step 1 — Request it as an optional feature
const session = await navigator.xr.requestSession('immersive-ar', {
requiredFeatures: ['local-floor'],
optionalFeatures: ['depth-sensing'],
depthSensing: {
usagePreference: ['gpu-optimized', 'cpu-optimized'],
dataFormatPreference: ['luminance-alpha', 'float32'],
},
});
Two things to get right here.
Put it in optionalFeatures, not requiredFeatures. A required feature that the device does not support makes requestSession reject, which means your whole experience fails to start on hardware that could have run it without occlusion. Request it optionally and check afterwards:
const hasDepth = session.enabledFeatures?.includes('depth-sensing');
The preference arrays are ordered, and they are preferences. The UA picks the first it can honour, so you must read back what you actually got rather than assuming your first choice:
console.log(session.depthUsage); // 'gpu-optimized' | 'cpu-optimized'
console.log(session.depthDataFormat); // 'luminance-alpha' | 'float32' | ...
Choose gpu-optimized for occlusion — you want the data as a texture, in the shader, never touching JavaScript. Choose cpu-optimized only if you need to read distances for logic: placing an object on a real surface, deciding whether a path is clear, triggering on proximity.
Step 2 — GPU path: get the texture
function onXRFrame(time, frame) {
const pose = frame.getViewerPose(refSpace);
if (!pose) return;
for (const view of pose.views) {
const depthInfo = glBinding.getDepthInformation(view);
if (!depthInfo) continue;
// depthInfo.texture — WebGLTexture
// depthInfo.width/height — its dimensions
// depthInfo.rawValueToMeters — scale factor
// depthInfo.normDepthBufferFromNormView — the alignment matrix
renderWithOcclusion(view, depthInfo);
}
}
glBinding is an XRWebGLBinding created over the session and GL context.
Two fields here are the ones people miss and then spend an evening on:
rawValueToMeters — the texture does not store metres. For luminance-alpha format it stores a 16-bit value split across two 8-bit channels, and you multiply by this factor to get metres. Skip it and your occlusion is wrong by a constant you cannot guess.
normDepthBufferFromNormView — the depth buffer is not aligned to your viewport. Its aspect ratio and orientation can differ, and it moves with device rotation. This matrix maps normalised view coordinates to normalised depth-buffer coordinates, and you must apply it when sampling. Occlusion that is right in the centre of the screen and drifts at the edges is almost always this matrix being ignored.
Step 3 — The shader
This is the actual occlusion. In the fragment shader for your virtual content:
precision highp float;
uniform sampler2D uDepthTexture;
uniform mat4 uDepthMatrix; // normDepthBufferFromNormView
uniform float uRawValueToMeters;
varying vec4 vClipPos; // the fragment's clip-space position
varying float vViewDepth; // distance from camera, in metres
float realDepthMeters(vec2 normViewCoord) {
vec2 d = (uDepthMatrix * vec4(normViewCoord, 0.0, 1.0)).xy;
vec2 packed = texture2D(uDepthTexture, d).ra; // luminance + alpha
float raw = packed.r + packed.g * 255.0;
return raw * uRawValueToMeters;
}
void main() {
vec2 ndc = vClipPos.xy / vClipPos.w;
vec2 normView = ndc * 0.5 + 0.5;
float real = realDepthMeters(normView);
// The real world is nearer than this fragment: the world wins.
if (real > 0.0 && real < vViewDepth) {
discard;
}
gl_FragColor = vec4(/* your shading */);
}
The whole mechanism is real < vViewDepth → discard. Everything else is getting the two numbers into the same units and the same coordinate space.
The real > 0.0 guard matters: zero means “no depth data for this ray.” The reconstruction has holes — dark surfaces, glass, mirrors, outside its range — and treating a hole as zero distance occludes everything. Fail open, not closed.
Step 4 — Soften the edge
The straight discard gives you a hard, aliased, low-resolution silhouette that looks like a cutout. Three fixes, in order of how much they buy:
Fade rather than cut. Replace the discard with an alpha ramp over a few centimetres:
float a = smoothstep(real - 0.05, real + 0.05, vViewDepth);
gl_FragColor = vec4(colour, 1.0 - a);
Sample more than once. Four or five taps around the point, averaged, hides most of the depth map’s low resolution for the cost of a few texture reads.
Be generous at silhouettes. Depth reconstruction is least reliable exactly at object boundaries, which is where occlusion is most visible. Biasing a centimetre or two toward showing your content produces fewer objectionable artefacts than biasing toward hiding it — a virtual object slightly overlapping a real edge reads as a near-miss; one with a chunk bitten out of it reads as broken.
Step 5 — CPU path, for logic rather than drawing
const depthInfo = frame.getDepthInformation(view);
if (depthInfo) {
const metres = depthInfo.getDepthInMeters(0.5, 0.5); // centre of the view
}
Normalised coordinates in, metres out, with the alignment handled for you. Useful for placement, for proximity triggers, for deciding whether to show a prompt. Do not use this for occlusion — pulling a buffer into JavaScript every frame and comparing per fragment is not going to hold frame rate.
Step 6 — The fallback, which you have to write
On iOS Safari and anywhere else without depth, you get nothing. The options, roughly in order of how well they work:
Plane detection. Floors and walls, which covers the most common and most jarring case — content sinking through the floor — without any depth map.
Hit testing for placement. Put content on real surfaces so that occlusion matters less.
Design around it. Content that sits in open space, above a table, or close to the viewer rarely needs occlusion. The effect is only missed when content is meant to be behind something.
Write the fallback first and treat occlusion as an enhancement. An experience that requires depth sensing excludes every iPhone, which for a web-delivered piece is most of your audience.
Where to go next
- Combine it with plane detection and anchors — depth for occlusion, planes for placement, anchors for stability. The three together are what make a scene feel sited rather than overlaid.
- Use depth for lighting. The depth map is a rough geometry of the room, which is enough for contact shadows where your object meets a real surface. Occlusion plus a contact shadow is a disproportionate improvement over either alone.
- Measure it. Sampling the depth texture several times per fragment is not free. Profile on the target device before adding the fifth tap.