Gaussian splatting scales badly in exactly the way that matters for the thing people most want to do with it. A room is fine. A building is manageable. A city block exhausts your VRAM long before you have captured anything like the whole scene, and the standard response — train it anyway, then prune — means you discover your memory problem after spending the compute.
Budgeted-GS: Real-Time Large-Scale Gaussian Splatting via Factoring LOD, posted 2 October 2026 by Haipeng Wang, proposes doing it the other way round.
The proposal
A factoring tree — described as “a multi-resolution hierarchy of moment-matched aggregates” — which allows flexible quality-to-memory tradeoffs and enables budget-centered training.
Three ideas in that sentence, and the third is the one that changes practice.
”Moment-matched aggregates” is the interesting technical choice
A level-of-detail scheme needs a way to replace many small things with one bigger thing that looks the same from a distance. For meshes you decimate triangles. For point clouds you cluster. For Gaussians the question is: given a cluster of small Gaussians, what single Gaussian best stands in for them?
Moment matching is the principled answer. A Gaussian is fully described by its first two moments — its mean (position) and its covariance (size, shape and orientation). So to aggregate a cluster, compute the mean and covariance of the combined distribution and build one Gaussian with those. The aggregate then has the same centre of mass and the same spatial spread as the group it replaces.
Why this is better than the obvious alternatives: averaging positions and taking the largest scale loses the group’s anisotropy — a cluster of splats lying along a wall should aggregate into a flat, wall-aligned Gaussian, and only the covariance carries that. Moment matching gets it for free, because covariance is the shape.
It is also the same mathematics as mipmapping, done properly. A mip level is a box-filtered average of the level below; a moment-matched aggregate is the distributional equivalent for an unstructured set of anisotropic blobs.
A multi-resolution hierarchy of these is then a tree where each node is a valid stand-in for its whole subtree, which is exactly the structure a renderer needs to make per-frame decisions about how much detail to draw where.
Budget-centered training is the part that inverts the workflow
The usual order is: train, then compress. You optimise a splat scene to minimise photometric error, end up with however many million Gaussians that takes, and then try to get it under your memory limit by pruning, quantising or distilling — all of which are lossy operations applied to a result that was not designed to survive them.
Budget-centered training means the budget is an input. You state the memory you have, and the optimisation allocates Gaussians within it — spending detail where it buys the most and aggregating where it does not.
That is a much better match for how work actually gets deployed. Nobody renders on an unspecified machine. You know whether the target is a Quest, a phone browser, a mid-range laptop or a workstation, and those are wildly different budgets. A method that takes the budget as a parameter produces one trained asset per target rather than one asset and a series of lossy derivatives.
And “flexible” tradeoffs implies the tree is traversable at runtime, which is the other half: a single trained scene that can be rendered at several budgets by descending more or less deeply, rather than several separately trained scenes.
Where this sits in a crowded month
This is the fourth splatting paper we have covered in two weeks, and together they describe a field working on the same problem from four sides:
- Stereo evaluation — splat artifacts are far worse in a headset than monoscopic metrics report
- Luminance dominance — geometry is built almost entirely from brightness, so chroma is compressible
- GALA — distil the expensive per-frame neural decode into a linear basis
- Budgeted-GS — make memory a training input rather than a post-hoc problem
The connective tissue is that splatting’s quality/cost frontier is being attacked at every stage of the pipeline simultaneously, which is what a technique looks like when it is moving from research into production.
The caution is also consistent across them, and the stereo paper is the one that supplies it: an efficiency gain measured monoscopically may be spending exactly the quality that stereo viewing needs. A LOD scheme that aggregates aggressively will produce larger, softer, more transparent Gaussians at distance — and soft semi-transparent blobs are precisely the artifact that stereo exposes and flat metrics forgive. Anyone adopting this for XR should check it in a headset before trusting the budget.