Emerging Interfaces

Designers Talk to Themselves While They Draw, and AI Tools Throw That Away

CommSketch takes the sketch and the concurrent speech together. Twenty-four participants reported significantly higher perceived alignment with the AI than sketching alone.

Sit near a designer sketching and you will hear them talking. Not to you — to the page. “So this is the body, and the handle comes round like… no, more like that. And this needs to be heavier.” It is continuous, half-audible, and it is where the actual thinking is.

Every sketch-based AI tool discards it. You draw, you press a button, and the model gets the marks.

CommSketch: How Speaking while Sketching Steers Human–AI Design Ideation, posted 29 September 2026 by Weiyan Shi, Darryl Lim, Geraldine Quek and Kenny Tsu Wei Choo, takes both.

The gap it identifies

Designers naturally verbalise their thinking while sketching, but most tools ignore this concurrent speech.

CommSketch is a sketch-based AI design interface that processes visual sketches and accompanying spoken commentary simultaneously, rather than as separate inputs.

The word doing the work is concurrent. This is not “describe your sketch afterwards”, and it is not “type a prompt then draw a mask”. It is the speech that happens while the hand is moving, which is a different kind of utterance entirely.

Why concurrent speech carries information a sketch cannot

A sketch is a record of decisions. The speech alongside it carries the things a sketch structurally cannot represent:

Intent versus accident. A wobbly line might be a deliberate organic curve or a failure to draw a straight one. The sketch is identical in both cases. “That should be straight” resolves it instantly, and no amount of model capability can infer it.

What matters and what doesn’t. Designers sketch at uniform fidelity because a pencil has one setting, but they do not care about every mark equally. “The proportions are the point, ignore the detailing” is a weighting that is invisible on the page.

Revision in flight. “No, more like a loop” is a correction to a mark that is already down. A sketch shows the final state of the paper; the speech shows the trajectory, including the rejected option — which tells a collaborator what you are avoiding as well as what you want.

Reference and analogy. “Like a kettle handle” imports a whole object’s worth of constraint in four words. Drawing that reference would take minutes.

And hierarchy. “This bit’s the important one” is structural information about a flat image.

This is the same fundamental point as the think-aloud protocol in usability research: concurrent verbalisation gives you access to reasoning that retrospective accounts reconstruct and distort. Designers were already doing the think-aloud. The tools just weren’t listening.

What the study found

Twenty-four participants, comparing concurrent speech plus sketching against sketches alone:

  • Concurrent speech supported natural expression of design intent and efficient visualisation
  • Speech fostered shared understanding, producing significantly higher perceived alignment between human and AI
  • Concurrent verbalisation let participants better direct AI contributions during ideation

“Perceived alignment” is the result, and the word “perceived” is load-bearing. The paper is not claiming the outputs were objectively better. It is claiming participants felt the system understood them — and that they could steer it.

Which is, for an ideation tool, arguably the more important measure. The failure mode of generative design tools is not bad output; it is output you cannot redirect. You get something adjacent to what you wanted, you try to nudge it, and the next generation is adjacent in a different direction. The felt experience is of negotiating with something that is not listening, and it is why so many of these tools get abandoned after the novelty passes.

A channel that carries intent, correction and emphasis in real time attacks that directly.

The practical objection, and it is serious

Speaking while working is not universally comfortable. In a shared studio, an open-plan office, a café, or a classroom, narrating your thinking out loud ranges from mildly awkward to socially impossible. The think-aloud protocol is used in labs partly for this reason.

So the honest reading is that this is a strong interaction for private working conditions and a weaker one for public ones, and a real product would need a text or gesture path to the same information. The paper’s own conclusion — that future tools should prioritise dynamic alignment, expanded multimodal communication options, and genuine human-AI co-creativity — explicitly points at “expanded options” rather than speech specifically, which suggests the authors see this too.

The second caution: speech recognition on half-muttered self-directed narration is harder than on dictation. It is quiet, fragmentary, trailing-off, full of um and no wait. Those disfluencies are not noise — “no wait” is a semantic signal — but they are exactly what ASR systems are tuned to discard.

What to take from it

Look for the channel your users are already using that your tool ignores. The information was always there. Sketching tools throw away speech; writing tools throw away pauses and deletions; 3D tools throw away the camera moves you made while deciding. Each of those is a free signal about intent that costs the user nothing to produce.

And measure alignment, not just output quality. For a creative tool, can I steer this? predicts adoption better than is the output good? The two are measured differently and only one of them is usually reported.

Worth reading next to the Engage-to-Unlock study we covered on 1 October. The two are in productive tension: that paper found value in withholding AI until the user has formed their own ideas; this one finds value in a richer channel into the AI from the start. Both are about keeping the human’s intent in charge, approached from opposite ends.