Emerging Interfaces

Instead of Asking an LLM to Act Like a Disabled User, Build an Inspectable Model and Simulate It

NeuroDivSim uses generative AI to construct task, interface and environment models you can review, then runs them through deterministic simulation against explicit cognitive configurations.

The fashionable way to use a language model in design research is to ask it to be a user. Prompt it into a persona, hand it your interface, read what it says. It is fast, it feels like research, and for anything involving disability or cognitive difference it is close to indefensible — you are reading a model’s statistical impression of how a group of people is written about, and calling it evidence.

NeuroDivSim, posted 8 October 2026, is an interactive tool that takes the opposite approach to the same problem.

What it does differently

Rather than using an LLM to act as a simulated user, NeuroDivSim uses generative AI to construct inspectable task, interface, and environment models from a usage scenario.

Then: after human review, those models are combined with explicit cognitive reference configurations and processed through deterministic simulation.

Three load-bearing words, each doing specific work.

Inspectable. The AI’s output is a model you can read, not a verdict you have to accept. If the task model has the wrong steps, or the interface model missed a control, you can see that and fix it. A persona’s narrative response offers nothing to check.

Explicit. The cognitive assumptions are configurations you set, not characteristics a model inferred. Working-memory capacity, processing speed, attention — whatever the reference configurations encode — are parameters you can read and argue with. That makes the resulting claim falsifiable in a way “I asked it to pretend to have ADHD” is not.

Deterministic. Same inputs, same outputs, every time. Which means a difference between two runs is caused by the thing you changed, and a result can be reproduced by someone else.

The experiment that becomes possible

The capability this enables is the point:

This enables designers to hold a modeled usage situation constant while varying cognitive assumptions and tracing their consequences to interaction steps and rule-based design recommendations.

That is a controlled experiment, in a domain where controlled experiments are extremely difficult to run.

The honest reason simulated users are attractive is that the alternative is hard. Recruiting participants with specific cognitive profiles is slow and expensive; asking a small number of people to represent a category is both methodologically weak and an imposition; and the variation between individuals within any such category is larger than the variation between categories. So teams either do proper studies rarely, or they guess.

Holding the situation fixed and varying one cognitive assumption at a time gives you something neither of those provides: which step of your interface is sensitive to which assumption. Not “is this accessible” but “step 4 fails when working memory is constrained, and step 7 fails under time pressure.” That is actionable in a way an overall verdict is not.

And tracing to specific interaction steps is what makes it usable. A recommendation attached to a step is a thing a designer can act on this afternoon. A score is not.

What it is not

The paper is careful, and so is the appropriate reading of it.

This is not a replacement for testing with actual people. A simulation’s conclusions are conclusions about the model, and the model is a simplification chosen by whoever built the reference configurations. What it can do is tell you where to look — narrow the space before you spend a participant’s time, and catch the obvious failures before anyone has to experience them.

The authors frame the whole thing as exploring simulation as an inspectable mechanism for reflecting on cognitive diversity during design and prototyping. Reflecting on, not determining. That is the right claim and a modest one.

The evaluation is correspondingly modest: an exploratory pilot with N=10, which provided formative insights into how participants engaged with the workflow and informed refinements to how models, results and recommendations are presented. Ten participants in a pilot is not evidence that the tool improves design outcomes, and the paper does not say it is.

Why this belongs in interface work generally

The pattern here extends well past accessibility, and it is the most transferable thing in the paper: use generative AI to build the model, not to be the subject.

Anywhere you are tempted to ask a model to simulate a person — users, audiences, reviewers, players — the same substitution is available. Have it construct an explicit, inspectable representation of the situation, review it yourself, and then reason over that representation with something deterministic. You trade a fluent answer for a checkable one, which is almost always the right trade when the answer is going to inform a decision.

Filed under cs.HC.