Ask a text-to-image model for an analog clock and it will very likely hand you one reading 10:10. Not because 10:10 is a natural time, but because watch advertising has set hands to 10:10 for decades — the V shape frames the brand name and looks cheerful — and the models learned from the photographs.
It’s Always 10:10: Reference Images Break a Bias That Prompts Only Dent, posted 8 October 2026 by Luca Cazzaniga, measures how strong the habit is and tests three ways of overriding it.
The design
52 models on the Magnific platform, with a replication on Higgsfield. Every image has to show three identical clocks reading 2:35, 6:50 and 11:20 — three different times, so a model cannot pass by getting lucky once.
The object description stays fixed; only the request about the time changes:
- (A) no time requested
- (B) the time in digits
- (C) the hand positions described in words, by construction, relative to the dial numerals
- (D) the same verbal description plus a drawn reference dial
1,799 images were read blind from coded copies by two AI readers, with a third reader and the author resolving disagreements. Dial-level agreement between readers: 96.0% and 97.3%.
What it found
With no time requested, 67% of images have all three clocks at 10:10. Two thirds of the time, unprompted, you get the advertising pose.
On the 20 current models, all three clocks are correct in:
| Condition | All three correct |
|---|---|
| (B) time in digits | 34% |
| (C) hand positions in words | 30% |
| (D) words + drawn reference dial | 75% |
The headline contrast is D−B: +37 points, 95% CI +28 to +45. And the one that should change how people write prompts: C−B: −4 points, CI −10 to +1 — no evidence that describing the hands in words beats simply writing the digits.
The replication on the 12 models shared by both platforms gives the same shape: B 54%, C 50%, D 81%.
Why the C result is the useful one
The instinctive response to a model getting a clock wrong is to explain harder. Describe the hour hand between the 2 and the 3. Describe the minute hand at the 7. Specify the angle. Everyone who has fought a diffusion model has done this, and the prevailing folk wisdom holds that more precise language produces more precise results.
This study measures that belief directly and finds it does not hold. Thirty per cent against thirty-four — the elaborate construction-based description performed no better than typing “2:35”.
That is worth sitting with, because it says something about where the failure actually lives. If more specific language did not help, the model is not misunderstanding the request. It understood “2:35” perfectly well and produced 10:10 anyway, because the prior over what a clock face looks like is doing the work, and the text conditioning is too weak to move it. You are not failing to communicate. You are being outvoted.
And why the reference image works
Supplying a drawn dial jumps full accuracy to three quarters and almost eliminates images entirely at 10:10.
The mechanism is not mysterious: an image conditions the generation in the same space the generation happens in. The prior over clock faces is visual, and a visual instruction competes with it on equal terms; a text instruction has to cross a modality gap first, and arrives weakened.
The practical rule generalises well beyond clocks. When a model keeps producing a canonical version of something instead of the specific version you asked for, stop rewriting the prompt and draw it. A rough sketch, a screenshot, a reference photo, a crude diagram in any paint program — a weak image beats a strong sentence. Hand positions, the number of fingers, the layout of a control panel, which way a mechanism faces, how many windows a building has: all of these are the same failure, and all of them are more tractable with a reference.
The corollary is also useful: 25% of images still had at least one wrong clock even with the reference. This is a large improvement, not a fix. If exactness matters, verify rather than assume.
The open-data part deserves a mention
The paper releases all images, prompts, raw readings and a script that recomputes every result — deposited at Zenodo with a DOI. Sixteen pages, five figures, six tables, filed under cs.CV with cs.HC.
Published benchmark claims about image models are typically a number in a table with no way to check it. A complete release of the stimuli and the readings means anyone can re-run the analysis, re-read the images with a different method, or extend the design to a different canonical pose. That is the difference between a result and a claim.