For about five years, building for VR meant picking a vendor. Oculus had a mobile SDK and a PC SDK. SteamVR had OpenVR. Windows Mixed Reality had its own. Each had its own initialisation, its own tracking model, its own input abstraction, and porting between them was a rewrite rather than a recompile.
OpenXR is the Khronos Group standard that ended that, and it is now implemented by essentially every runtime that matters: Meta Quest, SteamVR, Varjo, Pico, Vision Pro (via its OpenXR support), Monado (the open-source runtime), and the Android XR stack.
Khronos’ own tutorial session — long, but it is the authoritative introduction.
First: should you be using it directly?
Be honest about this before you write any code, because the answer is often no.
Use a game engine (Unity, Unreal, Godot) if you want to ship an application, you need physics, asset pipelines, audio, UI, and a scene graph, and you would rather solve your problem than solve rendering. All three support OpenXR as their backend, so you get the portability without touching the API.
Use OpenXR directly if you are building a custom renderer or a research prototype, you need precise control over the frame loop and latency, you are integrating XR into an existing C++ application, you are writing a tool or a runtime rather than an experience, or you specifically want to understand how this works.
Most creative-technology work belongs in the first category. This guide is for the second, and for anyone in the first who wants to know what their engine is doing.
The object model, which is where everyone gets confused
OpenXR has a hierarchy of handles and the names are not self-explanatory. Learn these five and the API becomes readable:
| Concept | What it is |
|---|---|
| Instance | Your connection to the OpenXR runtime. You create one, first, and it is where extensions get enabled. |
| System | A specific XR device configuration the runtime offers — in practice, “the headset”. You query for one by form factor. |
| Session | An active XR experience on that system. This is the thing with a lifecycle and states. |
| Space | A coordinate frame — view, local, stage, or one attached to a hand or controller. Everything positional is expressed relative to a space. |
| Action | A semantic input (“grab”, “teleport”), not a button. |
The critical mental shift: a session has states, and you must respond to them. IDLE, READY, SYNCHRONIZED, VISIBLE, FOCUSED, STOPPING. Your app is only receiving input and being seen by the user in FOCUSED. The runtime moves you between these — when the user opens a system menu you drop out of FOCUSED, and when they take the headset off you lose VISIBLE.
Handling that properly is the difference between an app that behaves and one that keeps rendering and taking input while the user is in the system settings.
The frame loop
This is the heart of it, and the ordering is not negotiable:
// Once per frame, in this order:
xrWaitFrame(session, &frameWaitInfo, &frameState);
// Blocks. The runtime decides when you should start rendering,
// to hit the display's timing. frameState.predictedDisplayTime is
// WHEN this frame will be shown — use it for all pose prediction.
xrBeginFrame(session, &frameBeginInfo);
if (frameState.shouldRender) {
xrLocateViews(...); // Where are the eyes AT predictedDisplayTime?
// ... render both eyes into swapchain images ...
}
xrEndFrame(session, &frameEndInfo); // Submit composition layers
xrWaitFrame blocking is the point, not a problem. The runtime knows the display’s refresh and its own compositor latency, so it tells you when to start work in order to finish just in time. This is how OpenXR keeps motion-to-photon latency low, and fighting it — by running your own timer — produces judder.
predictedDisplayTime is the only timestamp you should use. Not “now”. Every pose you request should be at the time the frame will be visible, because that is when the user’s head will be where the prediction says. Using the current time instead is the single most common source of “it feels laggy” in hand-rolled OpenXR.
Swapchains, not framebuffers. You acquire an image from a swapchain, render into it, release it, and reference it in a composition layer at xrEndFrame. The runtime’s compositor does the final work — distortion correction, chromatic aberration, timewarp/reprojection. You do not do lens distortion yourself; that was the old SDKs’ job and OpenXR took it away deliberately.
Action sets: the best idea in the API
This is where OpenXR genuinely improves on what came before, and it is worth understanding even if you only ever use an engine.
You do not read buttons. You declare actions.
// Declare what your app needs, semantically:
// "grab" — a float (trigger-like)
// "teleport" — a boolean
// "aim_pose" — a pose
Then you provide suggested bindings per controller type — Quest Touch, Valve Index, HTC Vive wand, a generic simple controller — mapping your actions onto that hardware’s inputs. At runtime the runtime picks the binding for whatever is actually connected.
Three consequences, all good:
- Hardware you have never heard of works, because the runtime maps your actions onto it
- Users can rebind, since the runtime owns the mapping (SteamVR exposes this directly)
- Your code reads as intent —
if (grabAction.currentState > 0.8f)rather thanif (trigger == A_BUTTON)
Always include a binding for the simple controller profile as a fallback. It costs five minutes and it is what makes your app run on hardware that did not exist when you shipped.
Spaces, briefly
Three reference spaces you will use:
VIEW— centred on the viewer, moves with the head. For head-locked content.LOCAL— fixed at the position where the session started. Seated experiences.STAGE— the room-scale play area, with its origin on the floor. Standing and room-scale.
Plus action spaces, created from a pose action, which is how you get controller and hand positions. Locate any space relative to any other with xrLocateSpace, at predictedDisplayTime.
Practical setup
The loader and headers come from OpenXR-SDK (Khronos, Apache 2.0). On desktop you link the loader and it finds the active runtime. On Android, the runtime ships in the vendor’s package.
For learning, use Monado. It is the open-source OpenXR runtime, it has a simulated headset driver, and it means you can develop and debug without putting a headset on every two minutes. This is a much better workflow than it sounds.
XR_EXT_debug_utils is the extension to enable first. It gives you readable error messages instead of enum codes, and OpenXR’s enum codes are not fun.
Khronos publishes an official tutorial at openxr-tutorial.com covering D3D11/12, OpenGL, OpenGL ES and Vulkan, with complete working code. It is the best written resource and it is free.
The honest assessment
OpenXR is verbose. A minimal “render a triangle in a headset” application is several hundred lines of initialisation — enumerate extensions, create instance, get system, enumerate view configurations, create session, create swapchains, create spaces, create actions, suggest bindings, attach action sets, then the frame loop. It is a Khronos API and it reads like one.
What you get is portability that actually holds, a frame loop designed by people who understand display timing, and an input model better than any vendor SDK produced. For a tool or a custom engine that needs to outlive a hardware generation, that is the correct trade.
For a piece you want to show next month, use Godot or Unity — and know that OpenXR is underneath.