Odyssey has opened a public research preview of Odyssey-3, a world model that generates interactive environments from text prompts and lets you move through them in real time. Founders Oliver Cameron and Jeff Hawke first showed it on 15 September; what is new is public access, technical detail, and benchmark results.
The free online demo runs Odyssey-3 Flash. You describe a world, choose first-person or third-person, move, trigger events, and watch the model respond.
A hands-on run through the public demo.
What it actually is
An autoregressive diffusion transformer that continuously generates new video frames from previous frames and user actions. The base model is 14 billion parameters generating at 832 × 480; Odyssey-3 Pro supports 1280 × 720.
The training mix is the part worth reading carefully, because it explains both what the model is good at and where it will fail:
- Internet videos with event descriptions
- Video game footage paired with the corresponding keyboard and mouse inputs
- Simulated physical interactions
The middle item is doing most of the work. Pairing frames with the inputs that produced them is what makes this controllable rather than merely generative — the model has learned a mapping from “W was pressed” to “the view moved forward,” from thousands of hours of someone playing. It also means the model’s notion of agency is shaped by how games respond to input, which is not how the world responds to input.
An additional training technique reduces the number of compute steps required, which is what makes real-time generation possible at all. Odyssey says the model learns physical relationships and cause-and-effect from visual observation during training.
Developers can apply for API access. The stated ambition goes beyond worlds to look at: Odyssey describes the model as a foundation for controlling systems — robotic arms and drones in simulation, characters in games.
The benchmark claim, and why it does not stand
Odyssey reports 66.1 points for Odyssey-3 Pro on the video-to-video benchmark from Physics-IQ Verified — a test that asks models to continue videos of real physical experiments and compares the continuation against what actually happened, across fluid mechanics, optics, solid mechanics, magnetism and thermodynamics.
That number is presented as a record, and it is not one.
The 66.1 figure comes from a single test run in which a selection method picked one of eight generated videos for each task. The benchmark’s own rules require four test runs with the standard deviation reported for a record claim. Without the selection method, Odyssey-3 Pro averaged 63.37 points across four runs.
Both numbers are on the official leaderboard, and both were submitted by Odyssey itself.
Three separate problems are stacked there, and they are worth naming because this pattern recurs constantly:
Best-of-eight is not a model score, it is a model-plus-selector score. If a selection method picks the best of eight generations, you have measured a system that includes a chooser — and unless the chooser ships with the model and runs in real time, that is not the thing you can use.
One run is not a measurement. Generative models have high variance across seeds. The benchmark requires four runs and a standard deviation precisely so that a lucky run cannot become a headline, and the rule exists because this keeps happening.
Self-submitted. Not disqualifying on its own — most leaderboard entries are — but it means nobody independent has reproduced either figure.
The honest version of the result is 63.37 across four runs, which is a real number and may well be a good one. It is also the number that is not in the announcement.
What it is actually for
Set the benchmark aside and the interesting question is what a real-time, promptable, explorable world is good for today.
Not finished work. At 832 × 480, frame-by-frame, with no persistence guarantees, this is not a renderer. What it generates exists while you are looking at it.
Possibly very good for sketching a space. The gap between “I want a long concrete corridor with light coming from the left” and something you can walk down is currently a day of modelling. If it is a minute, the activity changes — you try twelve corridors instead of committing to one.
And potentially useful as a behaviour, not an image. The framing Odyssey emphasises — a foundation for controlling robots and characters — points at the model as a predictor of consequences rather than a producer of pictures. A system that can answer “what happens next if I do this” is a different kind of component from one that draws.
The caution for anyone planning around it: a world model trained substantially on game footage has learned game physics, game affordances and game camera behaviour. Those are a specific and highly conventionalised subset of how things move and respond. That is an asset if you are making something game-shaped and a hidden constraint if you are not.