Digital Artists

An Agent That Plans a Subdivision Cage, Inspects Its Own Result, and Rolls Back When It Made Things Worse

The planner never generates a single vertex. It decides where control is needed and which curves carry design intent; geometry tools execute and verify every change.

A subdivision surface describes free-form geometry through a sparse control cage — a low-polygon mesh whose smoothed limit is the final shape. Recovering that cage from a dense scanned or sculpted mesh is the task every modeller knows as retopology, and the paper’s framing of why it is hard is exactly right:

it is not merely a fitting problem: the system must infer where control is needed, which curves encode design features, and when an initial result should be revised.

Those are three judgements, not three measurements. SubDGuide, posted 8 October 2026, is a modeler-inspired agentic workflow built around them.

The two stages

Stage A reads aligned multiview geometric evidence and produces a compact plan covering cage resolution, feature mapping and broad-form fitting. Not geometry — a plan.

Stage B inspects the resulting surface, requests targeted diagnostics, and chooses one of three actions: repair, rollback, or stop.

And the constraint that makes the whole thing credible:

Geometry tools execute and verify every change; the planner never generates vertices or connectivity.

Why that division is the contribution

This is the right architecture for applying a language-or-vision model to geometry, and it is worth being explicit about why, because the obvious approach fails badly.

A model asked to output a mesh will output a plausible mesh. Vertices in roughly sensible positions, connectivity that looks like connectivity, and — reliably — non-manifold edges, flipped normals, duplicate vertices, and quads that are not planar enough to subdivide cleanly. It has no mechanism for guaranteeing any geometric invariant, because it is predicting tokens that happen to be coordinates.

A model asked where control is needed is being asked the thing it is actually good at. “This crease is a design feature, that bump is scanning noise” is semantic judgement about what the object is. “Put a loop of edges along here” is a decision, not a computation. The computation — building the loop, fitting the surface, measuring the error — goes to geometry tools that cannot produce an invalid mesh.

So the planner supplies judgement and the tools supply correctness, and every change is verified rather than trusted. That is the pattern that separates agentic systems that work from ones that demo.

The numbers

On a fixed evaluation cohort, SubDGuide improves all five reported metrics over automatic remeshing:

  • Median Chamfer-L1: 0.641% → 0.443% of the target bounding-box diagonal
  • F-score at 1% tolerance: 82.00% → 92.78%

Two further findings matter more than the headline figures:

Stateful feedback improves four of five reported median metrics over one-shot planning. Stage B is doing real work — looking at the result and revising beats planning once and committing, which is the whole argument for the inspect-and-act loop.

Verification retains the earlier checkpoint when a proposal is unhelpful. The rollback is not decoration. The system proposes changes that make things worse, and catches them.

That last point is the one worth generalising. An agentic pipeline without rollback accumulates its own mistakes, because each step takes the previous step’s output as given. A verified checkpoint means the worst case is “no improvement” rather than “progressively corrupted.” Any iterative system acting on a document, a scene or a codebase has the same requirement.

The same interface supports six multimodal planners, which suggests the structure rather than any particular model is carrying the result.

What this would mean in practice

Retopology is one of the genuinely tedious bottlenecks in 3D work. A scan or a sculpt arrives as a dense triangle soup; getting to a cage you can animate, UV, edit and subdivide is hours of manual loop placement, and the existing automatic tools produce topology that is technically valid and useless to work with — uniform quads that ignore where the form actually needs control.

The distinction this paper draws is exactly the one those tools miss. Automatic remeshing optimises distribution; a modeller optimises editability. A cage with loops following the crease, the silhouette and the places you will need to push and pull is worth more than a cage with a better average error, and it is a semantic judgement about intent rather than a geometric one about surfaces.

The caveats to keep in view: this is evaluated on a fixed cohort with no stated breadth, the error metrics measure fidelity to the input rather than how pleasant the cage is to edit — which is the thing a modeller actually cares about and nobody has a metric for — and no code or availability is mentioned in the available material.

Filed under cs.GR.