XR / Spatial Computing

Put the AI's Explanation in the Room Instead of on a Second Screen

Situated explainability in AR means the saliency map sits on the actual object, in real time, rather than on a 2D dashboard you have to mentally re-project.

Explainable AI produces a lot of pictures. Saliency maps, attention heatmaps, feature attributions, counterfactual examples — almost all of them rendered as 2D images on a screen, derived from data captured earlier.

Which leaves the person a task nobody designed: look at an abstract heatmap, and work out which part of the real world it is talking about.

Seeing through the Eyes of AI: Situated Explainability in Augmented Reality, posted 2 October 2026, closes that gap. The authors are Ana Stanescu, Lucchas Ribeiro Skreinig, Tobias Langlotz, Stefanie Zollmann, Peter Mohr, Dieter Schmalstieg, Mark Billinghurst and Denis Kalkofen — a lineup that covers much of the serious AR research community.

The proposal

Rather than showing interpretability data on separate 2D screens using pre-recorded information, embed explainability directly into the user’s physical workspace through AR — delivering “spatial explainability information directly in a user’s workspace, in real time, as they explore the world.”

The stated contributions: that known explainability methods can be effectively adapted for AR, plus empirical findings on how users experience spatially integrated explanations.

Why the re-projection problem is worse than it sounds

Consider the actual cognitive task of reading a conventional saliency map.

A camera sees a scene. A model classifies it. A heatmap shows which pixels drove the classification. You are then shown that heatmap — on a laptop, flattened, from the camera’s viewpoint rather than yours, at a different time from when you were looking at the thing.

To use it you must: identify what the image is of, establish the camera’s pose relative to yours, locate the highlighted region in the image, map that region onto the physical object in front of you, and remember what the scene looked like at capture time.

Five mental operations before you learn anything. And every one is a place to get it wrong — which is why so much XAI research finds that users misread explanations, and why “explanation” and “understanding” come apart so reliably.

Painting the saliency onto the object removes all five. The highlight is there, on the thing, from your viewpoint, now.

Why real-time is the harder half of the claim

“Pre-recorded” is the other word doing work in the problem statement, and relaxing it is not free.

Most XAI methods are computationally expensive by construction. Gradient-based saliency needs a backward pass. LIME perturbs the input many times and fits a local surrogate. SHAP approximates Shapley values over feature subsets, which is combinatorial. Integrated Gradients accumulates gradients along a path, typically 50 to 300 steps. These are all designed for offline analysis, where a few seconds per explanation is irrelevant.

An AR headset needs the explanation registered to the world at frame rate, and it has a mobile GPU already busy doing passthrough, tracking and rendering. So “known explainability methods can be effectively adapted” is the interesting claim — the adaptation is presumably where the engineering is, and it is the part a creative technologist would want to read the paper for.

Why this matters past explainability

Situated information display is a general principle, and XAI is one instance of it.

The pattern — put the information on the thing it is about, from the viewer’s position, at the time it is true — is the argument for AR as a medium rather than as a novelty. And most AR content fails it. A floating panel of text in a headset is a screen that happens to be in the air; it has not become situated just because it is stereoscopic.

The things that do satisfy it are the AR applications that have actually stuck: surgical navigation, maintenance overlays, construction layout, speaker-attributed captions marked on the speaker’s face. In every case the win is the removal of a re-projection step.

For installation and interactive work the transferable question is: what is your audience currently having to mentally map? If a piece shows a visualisation of what its camera sees, or a graph of sensor data, or a readout of its own state, the audience is doing re-projection. Putting that information onto the physical thing — with projection mapping, with an LED on the object, with a material that displays its own state — is usually both clearer and more interesting than a screen.

And there is an honest caution. A situated explanation is more persuasive than a dashboard, because it appears to be a property of the world rather than an output of a model. That is the benefit and the hazard: an explanation painted convincingly onto an object will be believed more readily than it may deserve. Saliency maps are known to be unreliable and sometimes to pass sanity checks they should fail; rendering one in AR does not make it more faithful, only more credible.