XR / Spatial Computing

Generated 360° Video That Your Two Eyes Can Actually Agree On

EPIC extends video diffusion to 4K stereoscopic 360°, using an epipolar-aware matching metric as a preference signal — because stereo inconsistencies you'd shrug off on a screen are intolerable in a headset.

Immersive video has a content problem that generative models have not solved, because they have not been pointed at it. Video diffusion produces rectangular clips for flat screens. A headset wants 4K per eye, 360 degrees, in stereo, and simply upscaling and reprojecting a generated flat video gives you something that looks wrong in a way that is hard to describe and easy to feel.

EPIC: Epipolar-Consistent 360° Immersive Stereo Video Generation, posted 30 September 2026 by Debabrata Mandal, Dongdong Fu, Jonathon Miller, William Villareal, Xi Peng and Praneeth Chakravarthula, is a generative pipeline aimed at exactly that gap.

The core observation

Inconsistencies that viewers tolerate on regular screens become highly disruptive when viewed through an immersive headset.

This is the same finding, arriving from a different direction, as the Gaussian splatting study we covered on 1 October — where preference for the better reconstruction jumped from 58.4% monoscopically to 78.2% in stereo, and standard metrics failed to penalise the difference.

Two independent results, two weeks apart, saying the same thing: stereo viewing is a far harsher evaluator than a monitor, and the field’s metrics are calibrated for the monitor.

What epipolar consistency means

Two cameras looking at the same scene from different positions are not free to disagree arbitrarily. Given the geometry between them, a point in the left image must lie somewhere along a specific line in the right image — its epipolar line. That constraint is the foundation of all stereo vision, and your visual system enforces it ruthlessly.

When the constraint is violated — when a feature appears in the left eye somewhere that no camera geometry could put it in the right — the brain cannot fuse the two views. You get binocular rivalry: shimmer, instability, a surface that refuses to sit at a definite depth, and after a few minutes, discomfort.

A generative model has no notion of this. It produces two plausible images. Plausible-and-plausible is not the same as geometrically consistent, and the error is invisible on a flat screen, where you only ever see one of them.

So the authors built an epipolar-aware 360° image matching metric that captures temporal and stereo geometric inconsistencies across views — a measurement of how badly the two eyes disagree, extended across frames so that it also catches inconsistency that flickers over time.

The 360° part makes it harder than conventional stereo. On an equirectangular projection, epipolar lines are not straight — they are curves whose shape depends on where in the sphere you are, and the mathematics degenerates near the poles. Building a usable epipolar metric on a sphere is real work, and it is the paper’s technical core.

Using the metric as a preference signal

The pipeline extends existing video diffusion models and trains with direct preference optimization using limited training data, with the epipolar consistency metric as the preference signal.

This is the clever economic move. There is almost no 4K stereo 360° training data in the world — it is expensive to shoot, requires specialist rigs, and the public corpus is tiny. Training a model from scratch on it is not an option.

DPO sidesteps the data problem by needing only comparisons, and the comparisons here are generated automatically by the metric rather than collected from humans. The model produces candidates, the metric says which is more epipolar-consistent, and the model is nudged toward that. No annotation, no stereo rig, and you are fine-tuning rather than pre-training.

Which is a pattern worth noting generally: when you can write down a measurable property of a good output, you can turn it into preference data for free. The quality of the result is then bounded by how well your metric captures what matters — which is why building the metric first, as these authors did, is the right order of operations.

Why it matters for making things

Immersive content is the bottleneck, not immersive hardware. The headsets are fine. There is very little to watch in them, because 360° stereo capture is expensive and the production pipeline is specialist. The authors frame this as “a scalable path for bringing generative content to immersive displays, allowing diverse mixed reality experiences on demand”, and that framing is accurate about the actual constraint.

For artists, the interesting part is not the content-farm use. It is that generated 360° stereo becomes a material — a background, an environment, a dream sequence, a place that could not be filmed. The current alternative is building it in a game engine, which is a different skill set and a different look.

The caution is the usual one, and it is sharper here. Generated stereo that is mostly consistent will still cause discomfort in a way that generated flat video does not — an artifact you can ignore on a screen is a thing your vestibular system argues with. Test in the headset, with other people, for longer than a minute. A metric improving is not the same as a viewer being comfortable, which is precisely the gap the splatting study measured.