Generative video is good for eight seconds. Past that it falls apart in ways everyone who has tried it recognises but few papers name precisely: the character’s jacket changes colour, the room gains a window, the story stops going anywhere.
A Google Research post from 24 September 2026 names the failure modes and proposes an architecture against each.
The four failures, named
Precise vocabulary here is genuinely useful:
- Semantic drift — subtle shifts in character appearance or scenery across a sequence
- Cascading failures — an upstream error corrupting everything synthesised downstream
- Feature drift — entities unintentionally changing over time
- Content collapse — the narrative failing to progress meaningfully
The last is the one least discussed and most fatal. A model can maintain perfect visual consistency and still produce four minutes in which nothing happens — a character standing in a room, coherently, forever.
The architecture
A multi-agent orchestration layer built on Gemini and Veo, treating video generation as “a global optimisation and world-state tracking problem” rather than a sequence of independent clip generations. Four components:
- AI video co-director — uses multi-armed bandits to select creative configurations across strategy, narrative mode and aesthetic archetype
- CANVAS — maintains persistent visual memory of characters, locations and object states
- A²RD — generates long sequences through retrieve–synthesise–refine–update loops
- VQQA — uses vision-language model critiques to iteratively refine prompts
Results: a peak quality score of 81.4 on GenAD-Bench, with reported “significant continuity gains” across benchmarks, generating minutes-long videos with maintained visual consistency.
Why the world-state framing is the real idea
Strip the acronyms and the insight is that long-form video is a state-tracking problem, not a generation problem.
A diffusion model generating a clip has no notion that a character exists persistently. Each clip is conditioned on text and perhaps a previous frame, then generated fresh. Consistency emerges by luck and conditioning strength, which is why it decays with length — there’s nothing to decay from, because there was never a representation of “this character, who has a red jacket, who is in this kitchen.”
CANVAS is that representation. Once character and location state lives outside the generator, in a structure you can query and update, consistency becomes a retrieval problem rather than a coincidence. That’s the same move that made long-context language work tractable: stop asking the model to remember, and give it somewhere to look.
The multi-armed bandit choice for creative configuration is the odd, interesting one. Bandits solve explore-versus-exploit — try new options versus reuse what worked. Applying that to “strategy, narrative mode and aesthetic archetype” is treating creative direction as a search problem with a reward signal. Whether that’s a reasonable model of directing is a real question, and the honest answer is that it’s a reasonable model of directing to a benchmark.
What this means in practice
Two things, one encouraging and one worth caution.
It’s a pipeline result, not a model result. No new video model was trained. The gains come from orchestration around existing Gemini and Veo models — which means the approach is, in principle, reproducible with other models by anyone willing to build the state-tracking layer. The ideas transfer even though the implementation isn’t public.
“Minutes-long” is the honest ceiling. This is not feature-length coherence. It’s the difference between an unusable eight-second clip and a usable two-minute sequence — a meaningful jump for anyone doing animatics, previz, installation loops or short-form narrative, and not the end of the problem.
And note what’s being optimised: 81.4 on GenAD-Bench. Benchmark continuity is measurable; whether a sequence is worth watching is not, and no amount of world-state tracking addresses content collapse in the sense that matters — a story with nothing to say.