Research & Innovation

Thirty-Six Printable Cards for Deciding How to Evaluate an Explanation

Built from a review of 82 studies, sorted by hand, tested on five projects. The format is the contribution — a card-sorting deck is a better coordination tool than a framework paper.

There is a recurring failure in interdisciplinary projects that has nothing to do with anyone’s competence: the computer scientist, the designer and the social scientist all agree the system should be evaluated, and each means something different by it. One means accuracy of the explanation, one means whether users understand it, one means whether it changes behaviour. Nobody notices the disagreement until the results are in.

XAI Evaluation Cards: A Practical Method for Designing Human-Centred XAI Evaluations, posted 1 October 2026 by Kristýna Sirka Kacafírková, Ivania Donoso-Guzmán, Denis Parra, Katrien Verbert and An Jacobs, is a tool for having that argument on purpose and early.

What it is

A deck of 36 printable cards, derived from an updated evaluation framework built out of an analysis of 82 prior studies.

You use them with the card-sorting method: lay them out, and prioritise the evaluation dimensions that matter for your specific project.

Tested with 13 participants across five projects, the finding was that card sorting “can organise and streamline the design of the evaluation process, encouraging a more comprehensive and multidisciplinary assessment.”

The problem statement is precise: evaluating XAI from a human-centred approach requires choosing among numerous dimensions and measures, “often in an ad hoc and fragmented manner.”

Why the deck, and not a paper

There is already a substantial literature of XAI evaluation frameworks. The authors read 82 studies to build this one. So the contribution is not that the taxonomy did not exist — it is the format, and the format is doing something a framework paper cannot.

A card deck forces a finite set. Thirty-six cards is a number you can lay on a table and see at once. A framework paper’s taxonomy is a tree you navigate; a deck is a space you survey. People choose differently when they can see all the options simultaneously.

Sorting is a physical act that surfaces disagreement. When three people from different disciplines have to put the same card in the same pile, the disagreement happens immediately and visibly, with a specific object to point at. The alternative — each reading the framework and forming a private interpretation — defers the conflict to the results.

Prioritisation is forced by the medium. You cannot put all 36 cards in the “high priority” pile without it looking absurd. A written framework invites “we will consider all of these,” which is how evaluations become unfocused.

And it is a tool, not a reading. A deck can be used in a two-hour workshop by people who have not read the paper. That is a much lower barrier to adoption than a 20-page framework, and it means the method reaches the designers and domain experts on a project rather than only the researcher who found the citation.

This is a well-established tradition that deserves to be better known outside design research — IDEO’s Method Cards, the Tarot Cards of Tech, Oblique Strategies, Envisioning Cards for value-sensitive design, the Moral-IT deck. All of them take a body of knowledge that exists in text and convert it into a format that makes a group decide rather than read. The conversion is the work.

Why this matters past XAI

Most of what we cover is not explainable AI. But the pattern underneath is directly applicable, and it is the reason to read this.

Whenever a field has too many legitimate things to measure, evaluation quietly becomes whatever is convenient. We published two pieces this morning about exactly that in other domains — text-to-3D evaluation where the measurement protocol varies more than the generators, and a benchmark dispute between two decision models. In both cases the problem is not that nobody knows how to evaluate; it is that the choice of what to measure is made implicitly, by default, and then reported as if it were the only option.

For creative-technology work the parallel is sharp. How do you evaluate an installation? Dwell time, repeat visits, survey responses, observed behaviour, critical reception, how many people took a photo, whether it still works in month four. All legitimate, all measuring different things, and the one you pick is usually the one your venue already collects.

A forced-prioritisation exercise at the start of a project — whatever form it takes — is the intervention. You do not need these specific 36 cards. You need to have decided, before you build, which three things you will treat as evidence, and to have had the argument with your collaborators while it is still cheap.

And the deck is printable, which means this is usable by anyone reading the paper rather than requiring a licence or a tool.