Photogrammetry gives you a mesh that looks like a model of a place. A Gaussian Splatting capture, seen in a headset, looks like the place. The reflections are right, the foliage is right, the soft edges are right, and the thing your brain does when you put it on is categorically different from looking at a photoscan.
Getting there is a short pipeline with three places it goes wrong.
PlayCanvas on their splat rendering and WebXR over WebGPU.
Step 0 — Have a capture worth viewing
This walkthrough assumes you already have a trained .ply. If not, our Gaussian splatting tutorial covers producing one.
Two capture habits matter much more in VR than on a monitor, and are worth re-shooting for:
Shoot at head height, and shoot the floor. A capture taken from a comfortable standing position with the camera pointed slightly down leaves the floor beneath the viewer poorly reconstructed — and in a headset you look down. A visible hole where your feet are destroys the effect faster than any amount of blur elsewhere.
Cover the room, not the subject. Monitor viewing has a fixed arc of plausible camera positions. A headset user turns around. Anything you did not photograph from several angles will look like a smear when they do.
Step 1 — Pick a viewer
Three realistic options, and they differ in what you have to build.
PlayCanvas — their engine has first-class splat support and WebXR over WebGPU, and is the most capable if you want interaction, physics or UI alongside the splat. Their SuperSplat editor is also the standard tool for cleaning and compressing a capture, and it publishes directly to a viewer URL.
Three.js plus a splat library. More assembly, maximum control, and the right choice if the splat is one element of a scene you are already building in three.js.
A hosted viewer. Several services will host a capture and give you a link that enters VR. Zero work, least control, and your scene lives on someone else’s infrastructure.
For a first result, SuperSplat to a published URL takes about ten minutes.
Step 2 — Compress, because raw PLY is not deliverable
This is the step that decides whether anyone sees your work.
A trained splat scene is commonly tens to hundreds of megabytes as raw .ply, because every splat carries position, scale, rotation, opacity and spherical-harmonic colour coefficients as float32. Over mobile data, into a headset browser, that is a scene nobody waits for.
The formats to know:
.compressed.ply— PlayCanvas’s quantised PLY. Usually around 10× smaller, loads everywhere their tooling does.- SOG (Spatially Ordered Gaussians) — sorts splats spatially and stores attributes as image textures, so standard image compression does the work. Dramatically smaller again, and streamable.
- SPZ — Niantic’s compressed format, broadly supported, similar intent.
Also cut the spherical harmonics. SH coefficients are what make a splat view-dependent — the reason reflections change as you move — and they are the bulk of the file. Dropping from SH degree 3 to degree 1, or to 0, shrinks the scene enormously. For a diffuse interior the difference is barely visible. For anything with gloss, glass or water, the view-dependence is the effect, and cutting it makes the scene look painted.
Decide per capture. Do not apply one setting to everything.
Step 3 — Understand why stereo breaks naive sorting
Here is the thing that is specific to VR and catches everyone.
Gaussian splats are semi-transparent primitives, and alpha blending is order-dependent. Correct rendering requires drawing them sorted back-to-front from the camera. Every splat renderer therefore sorts the splat list by depth each frame — and a full sort of a million splats is expensive, so implementations sort approximately, sort on the GPU, or reuse the previous frame’s order when the camera barely moved.
In stereo there are two cameras. Sorting for the left eye produces an order that is slightly wrong for the right eye, because the depth ordering of nearby splats differs between viewpoints 64 mm apart.
The practical consequences:
- Sorting once per frame and using it for both eyes is fast, and produces subtle differences in blending between the eyes. The visual system is extremely sensitive to interocular mismatch, and this reads as shimmer, instability, or a vague discomfort people cannot name.
- Sorting per eye is correct and doubles the sorting cost.
Any mature splat renderer handles this; a hand-rolled one typically does not. If a scene looks fine on a monitor and feels subtly wrong in a headset, this is the first thing to check — close one eye, and if the discomfort goes away, it is the sort.
Step 4 — Scale and floor alignment
Splat training has no idea of real-world scale. The reconstruction is in arbitrary units, with an arbitrary orientation, and the origin wherever the solver put it.
On a monitor, nobody notices. In a headset, wrong scale is immediately and viscerally wrong — a room at 80% feels like a dollhouse, a room at 130% feels like a cathedral, and both make people feel strange without knowing why.
So:
Set the scale from something you measured. A door is about 2 m, a kitchen counter 0.9 m, a standard step 0.17 m. Measure one thing in the real space before you leave it.
Align the floor to y = 0 and make it level. Use a local-floor reference space and get the splat’s floor to actually coincide with it. A floor that is a few degrees off level is a tilted horizon, which is one of the more reliable ways to make someone ill.
SuperSplat has transform tools for both; do it there and bake it in, rather than correcting in code every load.
Step 5 — Hold the frame rate
The budget is brutal: two eyes at 72–120 Hz. Splat rendering is fill-rate heavy — every splat is a screen-space quad with blending, and in a headset the near ones cover enormous numbers of pixels.
The levers, in order of effect:
Splat count. The most direct control you have. Crop aggressively in SuperSplat — delete everything outside the region the viewer can reach, delete the floating garbage around the edges of the reconstruction, delete the sky. A million-splat scene that the viewer only ever sees a third of is two-thirds wasted.
Spherical harmonic degree. Lower degree is less per-splat shader work as well as a smaller file.
Resolution scale. Rendering at 0.8× and letting the compositor upscale is often an acceptable trade for a splat scene, because splats are already soft. This is far less damaging here than it would be on crisp geometry.
Avoid putting splats very close to the viewer. A splat 20 cm from the eye fills the view and costs an enormous amount of fill. Design the reachable area so the dense detail is at arm’s length or beyond.
Measure on the target device, not the desktop. A desktop GPU hides every one of these problems.
Step 6 — What to actually build around it
A splat on its own is a place you can stand in and nothing more, and that gets old in about ninety seconds. The things worth adding:
Teleport, constrained to the floor you aligned. Smooth locomotion through a static capture is unpleasant and unnecessary; teleport is comfortable and the scene does not move while you travel.
A boundary. Walk far enough and you exit the reconstruction, at which point the scene becomes a cloud of fragments viewed from outside. Fade to black, or physically block it.
Something that is not the splat. A captured room plus one authored element — an object you can pick up, a sound in a location, a text that appears — is the point at which it becomes a piece rather than a demo. The contrast between the photographic capture and the obviously synthetic addition is itself the material.
And note the shape of the technique’s current frontier: research this week on reusing the previous frame’s render instead of re-rendering the splat scene for small camera moves, and on making a static splat scene move at all.