Research & Innovation

φ-RIE Cuts Movable Objects Out of a Gaussian Splat and Fills In What Was Behind Them

Scanning a room gives you a photorealistic shell you can't touch. This pipeline couples asset extraction to background completion so the hole gets filled properly.

Scan a room with 3D Gaussian Splatting and you get something that looks extraordinary and does nothing. The chair is photorealistic from every angle you captured and completely inert. You cannot pick it up, because there is no “it” — the chair’s appearance is entangled with the background, its unseen geometry was never observed, and the floor it’s standing on doesn’t exist underneath it.

For anyone wanting to use a capture — in a game engine, an installation, a simulator — that’s the wall. φ-RIE: From Photorealistic Reconstruction to Interactive Environments, posted September 22, 2026, is a direct attack on it.

The authors are Runyi Yang, Deheng Zhang, Xiaoye Wang, Kanzhi Wu, Lei Sun, Ajad Chhatkuli, Kunyu Peng and Luc Van Gool.

The problem, stated precisely

The paper frames the gap as three entangled deficits:

  1. Object appearance may remain entangled with the background — there’s no clean boundary to cut along.
  2. Hidden object geometry is unobserved — you never saw the back of the chair or its underside.
  3. Occluded background content is unobserved — you never saw the floor and wall the chair was hiding.

Robot simulation, which is the paper’s stated target, needs objects that “move independently, make contact, and reveal previously occluded surroundings.” Every one of those requires solving all three.

The idea

φ-RIE is a Gaussian-native pipeline — it works in the splat representation rather than converting to meshes and back — that turns selected objects into movable simulator assets while preserving the rest of the reconstruction.

The key observation is the part worth remembering: asset construction and source removal should be coupled. One object identity should define both the movable asset you create and the scene content you remove and complete. Prior approaches treat those as separate operations, which is how you end up with an asset that doesn’t quite match the hole it left.

Architecturally, Scene Observation supplies shared evidence to Coupled Scene Construction, which produces registered assets plus completed background Gaussians, ready for simulator-driven rendering in an Interactive Environment.

Background on the broader workflow this sits inside: building photorealistic simulation environments from Gaussian splats.

Results

On 50 ScanNet++ scenes, evidence-based selection and registration retry raise matched F1 at 20mm from 0.336 to 0.383 at fixed retention. Further tests demonstrate asset executability and manipulation gains over a single-generator baseline.

Be clear-eyed about what that number is. 0.383 is not a solved problem — it’s a meaningful improvement on a hard geometric matching metric at a tight 20mm threshold, on real scanned scenes rather than synthetic ones. The honest summary is that this works better than what came before, not that captured rooms are now drop-in interactive environments.

Why a creative audience should care

The paper is written for robotics, and the evaluation is about manipulation. The capability generalises.

Every artist who has scanned a space and wanted to stage it has hit exactly this wall: the capture is gorgeous and rigid. Moving one object means either rebuilding it by hand or accepting a smeared hole where it used to be. That’s the difference between a splat as a backdrop and a splat as a set.

Coupled extraction-and-completion is the piece that makes the second thing possible. A pipeline that hands you the object as an asset and a plausible reconstruction of what it was standing in front of is what you need for virtual production, for installation pre-visualisation, and for any piece where the audience moves something in a scanned room and expects the world to still be there behind it.