Almost all Gaussian splatting work is evaluated the same way: render a held-out view, compare it to a photograph, compute PSNR, SSIM and LPIPS, put the numbers in a table. Almost all of it is looked at the same way too — on a flat monitor, one eye’s worth of image at a time.
And a very large fraction of the actual demand for splatting is VR, where neither of those holds.
“Beyond Monoscopic Viewing: A Study on 3D Gaussian Splatting Quality in VR”, posted 29 September 2026 by Shreyas Shivakumara, Gabriel Eilertsen and Karljohan Lundin Palmerius, measures the gap.
The result
A head-mounted display study comparing how viewers perceived reconstructions under different viewing conditions:
| Viewing condition | Preference for the higher-quality reconstruction |
|---|---|
| Monoscopic (standard single-perspective rendering) | 58.4% — weak |
| Stereoscopic VR | 78.2% — and consistent across all participants |
58.4% is barely above chance. Shown one image at a time on a screen, people could only just tell which reconstruction was better. Put the same two reconstructions in a headset and the preference is clear, strong, and — the detail that matters most — consistent across every participant.
Consistency is what makes this a finding rather than a trend. It is not that stereo viewing made people more opinionated on average; it is that stereo viewing surfaced something everyone could see.
And the metrics missed it
Image-quality metrics (PSNR, SSIM and LPIPS) and stereo-aware metrics (iSQoe and StereoQA) show only modest differences and fail to penalise the artifacts.
This is the harder half of the result. The failure of PSNR and SSIM is expected — they are per-pixel and per-window measures with no model of binocular vision. The failure of iSQoe and StereoQA, which are specifically stereo-aware, is the surprise.
The authors’ conclusion is unambiguous: monoscopic evaluation and standard image-quality metrics substantially underestimate perceptual artifacts in 3DGS reconstructions for VR.
Why stereo is so much harder on a splat, specifically
This is worth understanding mechanically, because it tells you which artifacts to look for.
A 3D Gaussian is a view-dependent object. Each splat is an anisotropic blob with spherical-harmonic colour, and the renderer projects it to screen space, sorts it by depth and alpha-blends. The appearance of a region depends on the viewing ray, by design.
Your two eyes are two different viewing rays, about 63mm apart. So each eye gets a slightly different projection, with potentially a different splat sort order and different spherical-harmonic evaluation.
That creates failure modes which do not exist monoscopically:
Binocular rivalry. If a region resolves differently in each eye — a splat visible to the left eye and occluded for the right, or a different blend order — the visual system cannot fuse it. The result is shimmer, instability, or a surface that refuses to sit at a definite depth. Monoscopically, each of those images on its own looks perfectly acceptable.
Depth from a blob. A Gaussian has no surface. Monoscopically your brain infers depth from occlusion, shading and parallax, and a soft blob is readily accepted as a slightly out-of-focus object. Stereoscopically your brain is given a disparity value and expects a surface at that depth. A semi-transparent cloud where a wall should be is a direct contradiction, and the fused percept is a wall made of fog.
Floaters become unmissable. The stray splats everyone with a splat pipeline recognises — the drifting artifacts in empty space — are easy to ignore in a flat render and impossible to ignore in stereo, because they acquire a definite, wrong depth and sit in front of things.
Edges reveal themselves. Monoscopically a soft silhouette reads as anti-aliasing. In stereo, the two eyes disagree about where the edge is, and the edge becomes a region of uncertainty rather than a boundary.
None of this registers in a metric that compares one rendered image to one photograph.
What to do about it
Evaluate in the headset. Every time. This is the entire practical takeaway. If your splat is destined for VR, a monitor-based review is not a cheaper approximation of the real review — it is measuring a different thing, and it will pass work that fails.
Build the review into the loop early. The expensive version of this lesson is capturing a scene, training, reviewing on a monitor, approving, building a whole piece around it, and discovering the problem at the install.
Treat floaters and transparency as the priority. Given that these are what stereo exposes, aggressive pruning, opacity regularisation and densification tuning are worth more for VR work than squeezing another decibel of PSNR. Optimising for the metric is, per this paper, optimising for the wrong objective.
Be sceptical of leaderboards for VR purposes. A method that wins on PSNR/SSIM/LPIPS has been selected for monoscopic fidelity. This result says that selection pressure does not reliably transfer, which means the best published method may not be the best method for you.
We have covered a steady run of splat-compression and splat-efficiency work — teacher-student compression, uncertainty-guided opacity, downstream-aware conditioning. This paper is a useful corrective to all of it: compression and efficiency gains measured monoscopically may be spending exactly the quality that stereo viewing needs.