Research & Innovation

LLM-Generated Interfaces Pass the Brief and Repeat Each Other, With Creative Range at 0.581 Against 0.902 for Humans

Appropriateness above 90% for every model and 99.2% for the best. Originality and range are the problem, and the gap to the human baseline is not small.

As language models saturate capability benchmarks, what they are bad at becomes harder to state precisely. Design Creativity Bench, posted 8 October 2026 by Aman Rusia, Abhijit Bhole, Prashank Gupta and Dipanjan Dey, measures one of those things carefully: diversity and appropriateness in UI designs.

The result is clean enough to be useful. Models satisfy briefs almost perfectly and produce substantially the same design when they do it.

Three measures, and the second is the clever one

Originality — distinctiveness among models given the same prompt. Do two different models answer the same brief differently?

Creative range — how much a single model’s designs change between two prompts for the same UI goal in different product domains. This is the measurement worth dwelling on. Give one model the same functional goal — say, a checkout flow — framed once for a luxury watch retailer and once for a municipal parking service. A designer’s two answers would barely resemble each other, because domain is most of what design responds to. Creative range asks whether the model’s do.

Appropriateness — the share of a brief’s acceptance criteria each design meets.

Splitting diversity from correctness this way is what makes the result legible, because the two trade against each other and a single “creativity” score would hide it.

The numbers

MeasureModelsHuman baseline
Originality (same prompt, different producers)0.592 (95% CI [0.582, 0.602])0.764 (95% CI [0.751, 0.778])
Creative range (same goal, different domains)0.581 (95% CI [0.567, 0.597])0.902 (95% CI [0.884, 0.919])
Appropriatenessabove 90% for every model; 99.2% best98.0%

The confidence intervals do not overlap on either diversity measure. These are not marginal differences.

And on appropriateness, the best model beats the human baseline — 99.2% against 98.0%. Models satisfy stated acceptance criteria more completely than people do.

Creative range is where the real failure is

0.581 against 0.902 is the widest gap in the table, and it is the one with the clearest practical meaning.

Human creative range at 0.902 says that when a designer is given the same functional goal in two different product domains, the two designs are almost entirely different artefacts. That is not decoration. Domain carries the conventions, the register, the expectations about density and tone, the regulatory furniture, the audience’s prior experience of similar products. Responding to it is most of the job.

A model at 0.581 is producing roughly the same design and changing the labels. It has read the functional requirement and largely ignored the context around it — which is consistent with how these systems work: the functional requirement is explicit in the prompt and the domain’s implications are not.

The authors’ own summary: the default output of LLMs, though generally appropriate, is substantially more repetitive than the human baseline. And they call for “strong measures to address the issue.”

Why appropriateness being near-perfect makes this worse, not better

It would be easy to read “above 90% appropriateness” as the good news. It is the mechanism of the problem.

A model that reliably satisfies every stated criterion and produces a generic result is exactly the output that passes review. Nothing is missing. Every box is ticked. The brief was met. There is no defect to point at, which is precisely why the repetition is hard to catch and hard to argue against — “it does what we asked” is true, and the thing that is wrong is not in the acceptance criteria.

This is the specific risk for anyone using these tools at the start of a design process rather than the end. A first draft that is competent and unremarkable is harder to reject than a bad one, and it anchors everything that follows.

The measure of originality is also a finding about convergence

0.592 for same-prompt pairs from different models — different companies, different training runs, different architectures — against 0.764 for human-model pairs.

Different models agree with each other more than they agree with a person. Whatever is producing the convergence is shared across the field, which points at the training distribution and the post-training conventions rather than at any one model’s limitations. You do not get out of this by switching provider.

Using it

The practical reading is narrow and actionable. Do not ask one model for one design. Ask for several, across deliberately varied framings, and prompt the domain explicitly rather than assuming the model will infer what the context implies — because the measurement says it does not. And treat a first output that satisfies the brief as a reason for suspicion rather than a reason to stop.

Filed cs.HC with cs.AI.