Everyone who captures scenes for Gaussian splatting develops folk wisdom about it: get even lighting, avoid blown highlights, shoot plenty of overlap, don’t change exposure mid-capture. The advice is good and the reasons are usually vague.
“What Builds the Scene? Luminance Dominates Geometry Formation in 3D Gaussian Splatting”, posted 30 September 2026 by Rezvan Joshaghani and Steven Cutchin, gives one of those rules a number.
The experiments and the numbers
Controlled experiments supervising training with different colour channels:
- Geometry learned from luminance alone supports held-out reconstruction 0.085 dB below RGB-trained geometry, on average
- Luminance-only supervision produced geometry 3.9–5.5 dB superior to chroma-only supervision during the geometry formation phase
- Removing chromatic information from trained models and re-fitting the frozen geometry cost minimal quality
Their conclusion, stated with appropriate care: geometry formation in standard 3DGS is “strongly luminance-dominated but not luminance-exclusive”, and much of the colour appearance can be recovered after the spatial structure has formed from brightness cues alone.
0.085 dB is nothing. PSNR differences below about 0.5 dB are generally imperceptible; 0.085 is deep in the noise. Throwing away all colour information during geometry training costs essentially nothing. Meanwhile training on colour without brightness is 4–5 dB worse, which is a large and obvious degradation.
Why this is what you would expect, in hindsight
The mechanism is the same one that makes every image and video codec work the way it does.
Structure lives in luminance. Edges, occlusion boundaries, texture, shading gradients, specular highlights — all the signals that let an optimiser work out where a surface is — are overwhelmingly carried in brightness. Convert a photograph to greyscale and you can still see everything that is where. Keep only the chroma and you get coloured mush with no discernible form.
The human visual system is built around this asymmetry: far more rod cells than cones, and much higher spatial acuity in luminance than in colour. Which is why JPEG subsamples chroma at 4:2:0 and nobody notices, and why that has been standard practice since 1992.
3DGS is solving a correspondence-and-geometry problem by gradient descent on photometric error. Of course it leans on the channel where the geometric information is.
Spherical harmonics explain the second half. Each Gaussian stores view-dependent colour as SH coefficients, which are separate parameters from its position, scale and rotation. So colour and geometry are not entangled in the representation — which is exactly why you can freeze the geometry and re-fit the colour with minimal loss. The paper’s result is partly a statement about the representation’s structure, not just about optimisation dynamics.
What to actually do differently
Prioritise exposure over colour accuracy when capturing. A clipped highlight or a crushed shadow destroys luminance information in that region, and luminance information is what builds the geometry. A slight white-balance error does not. So: bracket if you must, protect your highlights, and stop worrying about whether the colour checker was in frame.
Even, diffuse lighting matters more than you thought, for a specific reason. Hard shadows are strong luminance features that move with the light, not with the surface. If your capture has strong directional lighting, some of the luminance structure the optimiser is fitting is the shadow, not the geometry — and shadows get baked in as geometry. This has always been the advice; now there is a mechanism for why it matters so much.
Grayscale and low-light capture is more viable than assumed. If geometry only needs luminance, then monochrome sensors — which are more sensitive, have no Bayer filter and no demosaicing — are a legitimate capture path for the geometry pass, with colour supplied separately. The same logic extends to infrared, and to very low-light capture where colour noise dominates.
A two-stage pipeline is a real optimisation. Train geometry on luminance, then fit colour onto frozen geometry. The paper demonstrates that this costs little, and the first stage has a third of the per-Gaussian colour parameters to carry.
And it bears on the compression work. There has been a steady run of splat-compression papers — teacher-student distillation, uncertainty-guided opacity, downstream-aware conditioning. This result suggests where the compressible redundancy actually is: the chroma channels of the SH coefficients are doing much less structural work than the luminance, and could plausibly be quantised much harder.
Worth reading alongside the finding we covered on 1 October that splat artifacts are substantially worse in stereo than monoscopic metrics suggest. Together they point the same way: the metrics we optimise splats against are not measuring what viewers see, and the signals that build the geometry are not the ones we obsess over in capture.