Research & Innovation

A Benchmark That Asks Agents to Write the Code That Would Produce a Video

4DCodeBench evaluates inverse graphics on dynamic scenes — not reconstructing geometry, but recovering the program that generates it.

There are two ways to get a 3D scene out of a video, and they produce very different objects.

Reconstruction gives you geometry — a mesh, a point cloud, a radiance field, a set of Gaussians. It is faithful to what was captured and it is inert. You can look at it from new angles and you cannot meaningfully change it. Ask for the chair to be taller and there is nothing to adjust.

Inverse graphics in the procedural sense gives you the program that would make the scene. Parameters, loops, functions, a scene graph. It is less faithful and it is editable — the chair’s height is a number.

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes, posted 2 October 2026, evaluates the second kind, with time included.

What the task is

The name decodes cleanly:

  • 4D — three spatial dimensions plus time. Not a static scene but a dynamic one, with motion to be recovered as well as form.
  • Code — the output is a program, not a mesh.
  • Agents — the systems under test are agentic, meaning they presumably iterate: write code, render it, compare to the target, revise.
  • Bench — a standardised evaluation, which is the contribution.

So: given a dynamic scene, produce the code that generates it, including how it moves.

Why this is the harder and better problem

Because a program is a representation with structure, and structure is what you need to do anything.

Three concrete consequences for creative work:

Editability. A procedural scene has handles. Change a parameter and the scene changes coherently — not just the one vertex you grabbed, but everything that depends on it. This is the whole argument for procedural workflows and it is why Houdini exists.

Compactness. A few hundred lines of code can describe a scene that would take gigabytes as geometry. A procedural city is small; a scanned city is enormous.

Generalisation. A program that generates this scene, parameterised, generates a family of scenes. Recover the code for one building and you can produce a street.

And including time forces the model to recover behaviour rather than a sequence of states. A per-frame reconstruction of a swinging pendulum is forty poses. The program is one line with a period and an amplitude — and only the program knows it is a pendulum.

Why agentic iteration is the right shape for it

Inverse graphics has a property most code generation lacks: the output is directly checkable by rendering it.

Ask a model for code that sorts a list and verifying it requires tests. Ask for code that produces a scene and you can render it and compare pixels. That gives you a dense, automatic, differentiable-ish reward signal with no human in the loop — which is exactly the condition under which agentic loops work well. Write, render, diff, revise.

It is the same structural insight as the EPIC paper’s use of an epipolar metric as a preference signal: when you can measure goodness automatically, you can generate the training or search signal for free.

The thing to watch for when the results arrive

Inverse graphics benchmarks are extremely sensitive to the vocabulary you allow, and this is where claims in this area usually need scrutiny.

If the target scenes were generated by programs drawn from a known library of primitives, then recovering the code is a search over a constrained space — hard, tractable, and not the same problem as recovering code for arbitrary real video. If the targets are real captured footage, the task is enormously harder and the scores will be correspondingly low.

Both are legitimate benchmarks. They measure different things, and a headline number means nothing without knowing which.

And a second caution, which this month keeps supplying: we published a piece two days ago on text-to-3D evaluation where the measurement protocol varied more than the generators did — 17 of 19 evaluators. A benchmark that scores code by rendering it inherits every rendering-configuration variable that paper identified: camera placement, lighting, resolution, tone mapping. The existence of a standard benchmark is progress; the assumption that its numbers are protocol-independent is the error that paper documents.

Worth watching all the same. Scene-as-code is the representation creative tools actually want, and it has been the missing piece between capture and authoring for a long time.