Leaderboards are how this field decides what to use. A new text-to-3D method appears, posts a table, tops a metric, and gets adopted. “Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation”, posted 30 September 2026 by Anson Y. Lam, Shuqing Li and Michael R. Lyu, asks whether the tables mean anything.
The experiment
The design is clean and the clue is in the title: freeze the scenes.
Hold the generated 3D content completely constant, and vary only the measurement configuration — the rendering settings used to produce images from the scene, and the caption wording used to test text-alignment. Then see whether the scores move, and whether the rankings move.
They do.
Peak configuration variance exceeds between-generator variance for 17/19 alignment evaluators.
Read that carefully. For seventeen of nineteen evaluation metrics, how you measured produced a bigger spread in scores than which generator you were measuring did. The measurement noise is larger than the signal it is supposed to detect.
And because the spread reshuffles the order, the winners shift — the same set of frozen scenes produces different rankings under different, individually reasonable, protocols.
Why this is so easy to get wrong
Both knobs they turned are ones nobody thinks of as a knob.
Rendering settings. A text-to-3D evaluation almost always works by rendering the 3D asset to images and scoring those images with a 2D model — typically a CLIP-style text-image alignment score. But rendering involves a hundred choices: camera distance and elevation, field of view, how many views, background colour, lighting, whether there is any lighting, tone mapping, resolution, anti-aliasing. None of these are part of the generated asset, and all of them change the pixels the scorer sees.
A grey background versus a white one changes a CLIP score. So does zooming in. Neither has anything to do with the quality of the 3D model.
Caption wording. The alignment metric compares the render against text. Change “a red sports car” to “a sports car, red” to “a crimson automobile” and the embedding moves. The asset has not changed; the question has been rephrased.
So the evaluation is measuring a composite of the asset, the renderer, and the prompt phrasing — and it is reporting it as a property of the generator.
The recommendation, which is practical
The authors propose reporting results as (generator, score, card ID) with protocol-dependent comparisons — that is, a score is meaningless without an identifier for the exact configuration that produced it, and comparisons are only valid within a configuration.
That is a modest, implementable fix, and it mirrors what other fields had to learn. It is essentially a model card for the measurement, and it is the same move as pre-registration in psychology or random_state in a scikit-learn paper: make the protocol a first-class part of the result rather than an unstated default.
The harder implication, which they state: leaderboard rankings lack robustness. Not that they are wrong, but that a ranking produced under one protocol does not transfer.
Why this matters today specifically
We are publishing two other pieces this round that depend on benchmark numbers, which makes this a useful corrective to both.
The decision-model piece reports Cloudflare’s Clef winning 6 of 10 on the Jev Decision Index. That suite was built by a competitor, which is a genuine strength — but independently constructed is not the same as configuration-robust, and the margins on several of those benchmarks are a couple of points. This paper is about text-to-3D, not typed decisions, but there is no reason the mechanism would be absent elsewhere: any evaluation involving a rendering step, a prompt template, or a scoring model inherits the same fragility.
And the splatting paper reports geometry differences of 0.085 dB and 3.9–5.5 dB. The second is clearly real; the first is explicitly described as negligible, which is the correct way to report a small number.
The general discipline this suggests, for reading any benchmark table in this field:
- Ask what the protocol was, and whether it was the same for every row
- Treat gaps smaller than a few percent as unresolved unless the paper reports variance across configurations
- Be most suspicious where a 2D model scores a 3D thing, because the rendering step is an unreported variable
- Test on your own material, in your own configuration, before adopting
That last one is not a counsel of despair. It is the actual point: a leaderboard is a shortlist, not a verdict.
Related Reading
- Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation — arXiv:2610.00447
- CLIP: Learning Transferable Visual Models From Natural Language Supervision — arXiv:2103.00020
- Reproducibility — Wikipedia
- Model cards for model reporting — arXiv:1810.03993
- ML Reproducibility Challenge