Everyone who has tried to upscale video with an image model has seen the same failure. The individual frames look sharper. Played back, the result crawls — skin texture shifts, hair reorganises, facial features subtly morph from frame to frame.
That’s not a quality problem. It’s a structural one: image upscalers process each frame independently and have no idea the other frames exist.
SeedVR2, from ByteDance Seed, is built for video, and it’s now natively supported in ComfyUI (PR #14424). It entered our tracker this week at 471,279 downloads.
What it is
A one-step diffusion-based video restoration model, adversarially trained against real data.
The “one-step” part is the engineering achievement. Conventional diffusion needs multiple denoising passes per frame; for video that multiplies into an unusable amount of compute. SeedVR2 does a single forward pass for both image and video upscaling.
The approach is described as conservative upscaling — preserving original structure and detail while enhancing clarity, rather than inventing detail that wasn’t there. For restoration work that’s the right bias: a model that hallucinates plausible texture is a model that changes what your footage shows.
The temporal consistency comparison
A concrete test that circulated with the ComfyUI integration: 540p talking-head footage upscaled to 1080p.
- ESRGAN — facial features visibly morphed and flickered
- SeedVR2 — facial features stayed stable, adding consistent texture to skin, hair and clothing that remained coherent across all 240 frames
A VFX compositor’s wider test across seven ComfyUI upscaling methods, five content types (fine detail, structures, defocused night, people, abstract, 2D anime) and three target resolutions found SeedVR2 the winner for versatility — balanced across content types — and best on 2D/anime, with FlashVSR faster.
The honest conclusion from that test is worth repeating: no single upscaler wins everything, and you should choose by content type rather than by leaderboard.
Model sizes and practical workflow
Two variants with various quantizations:
- 3B — lower VRAM, suitable for limited hardware
- 7B — higher quality, more VRAM and processing time
The community-recommended pattern is asymmetric and sensible: images on the large model (7B-FP8) with high step counts for maximum detail; video on the small model (3B-FP8) with lower weights for speed while keeping temporal consistency.
For 4K, the recommended approach is two passes — go to 1080p first to get a clean intermediate, then upscale that to 4K. One giant jump produces worse results than two controlled ones, which is the same lesson every traditional upscaling pipeline learned years ago.
Where this actually gets used
Not, mostly, in making AI video look better — though it does that. The more valuable applications are archival and production:
- Old footage — home video, tape, early digital, where the source is genuinely low-resolution rather than artificially degraded
- Documentary and archive work, where inventing detail is an ethical problem and “conservative” upscaling is the requirement
- Rescuing compressed footage that arrived as a low-bitrate export with no master available
- Installation content that needs to run at wall scale from a source that wasn’t shot for it
Conservative, temporally consistent restoration is a more useful tool than aggressive enhancement for all four. The frame-by-frame upscalers were never going to work for any of them.