XR / Spatial Computing

Stop Re-Rendering the Splat Scene for Every Small Camera Move and Reuse the Last Frame Instead

Warp the previous RGB-D render with the relative pose, then let a lightweight network predict the residual that fills in what the warp could not know.

Walking through a Gaussian Splatting scene means a new render every frame. But small camera motions preserve most of the visible content — move 2 cm and 95% of what you can see, you could already see. Conventional 3DGS renders the whole target view anyway, and that overlap is thrown away.

LVS: Local View Synthesis from Relative Camera Pose by Reusing Previous Views, posted 8 October 2026, spends that overlap instead.

The two halves, and why neither works alone

Geometric warping uses depth and relative pose to transport content from the source image to the target. This is standard, cheap, and has two well-known failure modes that the paper names precisely:

It cannot recover newly exposed content. Move the camera and the far side of an object comes into view. That content was occluded in the source frame, so there is nothing to warp — you get a hole, and the hole lands exactly at silhouette edges where it is most visible.

It is sensitive to depth errors. Warping transports each pixel to where the depth says it should go. Wrong depth puts it in the wrong place, and the error grows with the size of the camera move. Splat-derived depth is not exact, particularly on thin structures and semi-transparent regions.

So LVS adds the second half: a lightweight multiscale network predicts RGB residuals to correct artifacts and infer missing appearance. The warp does the heavy transport, the network fixes what the warp cannot know about.

On GS-render, residual refinement improves PSNR by 0.72 dB over pure warping.

Residual prediction is the right division of labour

The architectural choice here is worth separating from the result, because it is the reason the network can be small.

A network asked to synthesise the target view has to produce everything — the geometry, the textures, the lighting, all of it — and must therefore be large, because it is doing rendering.

A network asked to predict the residual on top of a warp inherits a nearly-correct image and only has to supply the difference: fill the disocclusion holes, clean up the stretching at depth discontinuities, fix where the depth was wrong. That is a narrow, local, low-energy signal, and it is learnable by something small and fast.

This is the same reasoning behind residual connections generally and behind temporal upscalers in real-time graphics — give the model a good starting point and ask only for the correction. The paper also caches source features to reduce repeated computation, which is the same economy applied to the network itself: the source frame does not change between nearby target views, so neither do its features.

Note that this is a per-scene framework. The network is trained for the scene it serves, which is consistent with how splat scenes are already handled — a splat is itself a per-scene artefact — but means there is a training step per capture, not a general-purpose model you point at anything.

Why this is an XR result rather than a rendering curiosity

The paper’s own closing line puts it in the right place: the separation of scene rendering from local view updates supports responsive scene exploration, with potential applications in augmented and virtual reality. The evaluations demonstrate low query latency on both captured and rendered scenes.

In a headset the constraint is not average frame rate, it is latency between head motion and the matching photons. Miss it and the world lags behind the head, which is both immediately noticeable and nauseating. The standard mitigation is reprojection — take the last rendered frame and warp it to the current head pose — which is exactly the first half of LVS, shipped in every headset runtime for a decade, with the same disocclusion artefacts at silhouette edges that everyone has learned to ignore.

LVS is, in effect, reprojection with a learned residual applied to splat scenes. If the residual network runs inside the latency budget, the disocclusion holes that reprojection has always left behind get filled by something that knows what the scene looks like. That is a meaningful improvement to a technique that is already load-bearing.

The two things to establish before planning around it: “low query latency” is not a millisecond figure, and a headset budget is unforgiving; and the method is evaluated for nearby views, with no stated bound on how far a camera can move before the warp-plus-residual approach loses to simply rendering. That threshold is the whole engineering question, and it is not in the abstract.

Filed cs.CV.