Research & Innovation

Users Said Half the AI's Reasoning Steps Were Things They Never Asked For

24 participants, nearly 1,000 annotations at the individual reasoning-step level — and about half of the steps were judged AI-initiated rather than delegated.

Ask a reasoning model to help with something and it does not just answer. It decides what the question means, picks an approach, makes assumptions, chooses what to include, settles on a format, and resolves ambiguities you had not noticed were ambiguous — and then presents the result as though you had specified all of it.

What Did the AI Take On? Characterizing Cognitive Delegation in LLM Reasoning, posted 5 October 2026 by Yoonsu Kim, Sean Kim, Kihoon Son, Saelyne Yang and Juho Kim, measures how much of that happens.

The finding

24 LLM users on knowledge-work tasks, with nearly 1,000 annotations at the individual reasoning-step level.

Users viewed approximately half of all reasoning steps as “AI-initiated,” meaning the model assumed responsibilities the users hadn’t explicitly requested.

Half.

And the tension the paper identifies is the one that makes this hard rather than merely interesting: important decisions might escape user attention when delegated, yet manually reviewing every step becomes impractical.

Both halves of that are true at once, which is why “just read the chain of thought” is not an answer. A twenty-step reasoning trace takes longer to audit than the task would have taken to do, and the whole point of delegating was to not do that.

Why delegation and initiation are genuinely different

Worth separating carefully, because the distinction is what the paper is built on.

Delegation is handing over a step you chose to hand over. Draft this section; I’ll edit it. You know the boundary, you know what to check, and you remain the author of the decision to delegate.

Initiation is the model taking a step you did not assign. You asked it to summarise a document and it also decided which points mattered, imposed a structure, and silently dropped a caveat. Nobody did anything wrong — the task was underspecified, and the model had to resolve the underspecification to proceed.

The problem is that initiation is invisible by construction. A delegated step has a boundary you drew, so you know where to look. An initiated step happened inside a gap you did not know existed, and you only discover it if you reconstruct the model’s reasoning and notice a choice you would have made differently.

Which you will not do, at least not often, because reviewing is the expensive thing you were avoiding.

The nuance that makes it a design problem rather than a complaint

The paper’s most useful result is not the 50% figure. It is this:

Desired user involvement fluctuated based on the type of cognitive work and how the delegation unfolded, even when the AI’s contributions matched user intentions.

Read that last clause carefully. Even when the model did what the user wanted, the user’s preference about being involved varied.

So this is not about accuracy. A model can take a decision, take it correctly, produce exactly the output the user would have chosen — and the user can still, legitimately, have wanted to make that decision themselves. Agency is a separate good from correctness.

That is why the authors’ proposed direction is flexible protocols and inspectable, revisable AI-initiated decisions rather than “ask permission more” or “be more accurate.” Different kinds of cognitive work want different involvement, from the same person, on the same day.

They built taxonomies of the kinds of cognitive work LLMs do, how delegation actually occurred, and what protocols users preferred — per reasoning step. Those taxonomies are likely the lasting contribution.

What this means for creative tools

This is the third study in a fortnight converging on the same territory from different angles, and together they are starting to say something.

Engage-to-Unlock found that withholding the model until the user had engaged with the task improved evaluation efficiency — keep the human in front of the AI.

CommSketch found that capturing concurrent speech while sketching significantly raised perceived alignment — give the human a richer channel into the AI.

And this paper locates the gap both are addressing: the model is making about half the decisions and the user cannot see which ones.

The practical design implication for anyone building a creative tool with a model in it:

Make the initiated decisions visible, not the reasoning. A full chain of thought is unreadable. A short list of here are four choices I made that you did not specify is reviewable in seconds and is where all the risk is.

Make them revisable in place. An initiated decision you can see but not change is just a disclosure. The value is in being able to override one of four assumptions without re-running the whole task.

And let the protocol vary. Some tasks the user wants to own; some they want gone. A fixed level of consultation will be wrong in both directions. This is the same argument as GEV’s confidence-gated escalation — adapt the amount of process to the situation rather than fixing it.