Emerging Interfaces

What Your Headset Screenshots Is Not What You Can See, and That Gap Is a Security Hole

A rectangular capture frame versus a non-rectangular human visible field — with prompt injection and privacy leakage among the consequences.

When an XR system needs to know what you’re looking at — to answer a question about your surroundings, to feed a vision-language model, to log context — it takes a screenshot. That frame stands in for your first-person view.

A paper posted September 24, 2026 points out that the substitution is invalid, and that the gap has consequences beyond inaccuracy.

“Through Human Eyes and Machine Eyes: Understanding View Mismatch in Video See-Through Extended Reality”, by Yanming Xiu, formalises the problem and demonstrates why it matters.

The mismatch

A screenshot records a rectangular, machine-readable frame. Your effective visible region through a video see-through headset is more constrained and non-rectangular — shaped by the optics, the field of view, and where your eyes actually are.

The paper defines the relationship precisely by naming three regions:

  • Co-visible — both the system and the human can see it
  • System-only — captured in the frame, but not visible to the user
  • Human-only — visible to the user, but not in the captured frame

A pilot-level boundary measurement on Meta Quest 3 confirms a clear mismatch between the rectangular screenshot frame and the approximate human-visible boundary.

The four failure modes

This is where it stops being a geometry curiosity. Through four case studies, the paper illustrates:

1. Prompt injection. This is the serious one. If there is a system-only region — pixels the camera captures that the wearer cannot see — then text placed there reaches the vision-language model without ever reaching the human. An adversary can put instructions in your headset’s visual field where you will never notice them. You cannot review what you cannot see.

2. Privacy leakage. The captured frame includes content outside what you believed you were sharing. If you screenshot to ask an assistant about the object in front of you, the frame may carry things at the periphery that you had no idea were in shot.

3. Human-invisible information bias. The model’s answer is influenced by content in the system-only region. You get a response shaped by evidence you never saw, with no way to reconcile it against your own view.

4. Missing human-visible information. The inverse: things you can plainly see are outside the frame, so the model doesn’t know about them. Its answer omits what is obvious to you, which reads as the system being stupid rather than blind.

Why this pairs with the passthrough research

We covered “Passthrough Rigidity” two days ago — a 110-person study finding that video passthrough suppresses head rotation four-fold as users adopt motor caution.

Put the two together and a picture emerges about the current generation of see-through XR: the human and the machine are having meaningfully different visual experiences of the same room. The user is moving less than they naturally would and seeing a constrained, non-rectangular region; the system is capturing a rectangle that includes things the user cannot see and excludes things they can.

Every assistant feature built on “the headset can see what you see” inherits that gap.

What to do about it

For anyone building on headset capture, the paper’s framing gives you the vocabulary to design against:

  • Don’t treat a screenshot as consent. If your system sends captured frames anywhere, the user has not reviewed the periphery.
  • Treat the system-only region as untrusted input. Text in a captured frame is not text the user approved, and should never be followed as instruction. That’s a general lesson from prompt injection, arriving here in a new form.
  • Crop toward the human-visible region when the point is to represent the user’s view — accepting that you’ll lose some field — rather than using the full sensor frame because it’s what the API returns.

The research is pilot-level and explicit about it; the boundary measurement is one headset and a small sample. The mechanism it names, though, is structural rather than incidental, and it will apply to every video see-through device until the capture frame and the visible field are made to coincide.