AI & Creative Tools

AI Upscaling, Honestly: What It Can Add and What It Invents

Which model for which job, why a 'restoration' upscaler will change your subject's face, and the one setting that fixes most bad results.

“Upscale this” sounds like one operation. It is two, they are in tension, and almost every disappointing result comes from using one when you wanted the other.

pixaroma’s walkthrough of the ComfyUI upscaling nodes, which is where most of this is practical to try.

The two jobs

Faithful enlargement. Make the image bigger without inventing anything. The output should contain exactly the information the input had, resampled. This is what you want for documentation, archival work, print preparation, and anything where the image is evidence.

Generative detail synthesis. Make the image bigger and hallucinate the detail that a higher-resolution capture would have had. Pores, fabric weave, individual leaves, text. This is what you want for stylised work, concept art, upscaling AI-generated images, and anything where the image is an illustration.

The second job is not recovering lost information. The information is gone. The model is making up something statistically plausible, and plausible is not the same as true.

This is why upscaled faces look like strangers. A generative upscaler given a 200-pixel face does not have enough information to reconstruct that person’s features, so it synthesises a face consistent with the blur. It will be sharp, detailed, convincing, and not them. If you are upscaling a photograph of a real person — family archive, documentary, journalism — this is a serious problem and not a quality setting.

The tools, by job

Faithful-ish: ESRGAN-family models

Real-ESRGAN is still the sensible default for 2x–4x enlargement that mostly respects the input. It is a GAN, so it does add detail, but it is trained to add texture-plausible detail rather than to reimagine content. Fast, runs on modest hardware, available everywhere.

Model variants matter more than people realise:

  • RealESRGAN_x4plus — general photographic content
  • RealESRGAN_x4plus_anime_6B — line art and anime. Using the photo model on line art produces mush; using the anime model on photos produces plastic.
  • realesr-general-x4v3 — a good compromise with a denoise strength control

4x-UltraSharp and 4x-NMKD-Siax are popular community ESRGAN checkpoints, and worth trying — ESRGAN architecture models are interchangeable, so you can swap checkpoints freely in the same node.

For genuinely non-generative enlargement, Lanczos or bicubic resampling is still correct, and you should not be embarrassed about it. For print, a clean Lanczos upscale with a light unsharp mask is frequently better than anything a model does, because it adds nothing to argue about.

Generative: diffusion-based upscalers

SUPIR, Stable Diffusion upscalers, and tiled diffusion approaches use a full diffusion model to regenerate the image at higher resolution, conditioned on the original. The detail they produce is far richer, and far more invented.

The key control is denoise strength (sometimes “creativity” or “restoration”):

  • 0.1–0.25 — detail enhancement, structure preserved. Start here.
  • 0.3–0.45 — substantial reinterpretation. Faces will change.
  • 0.5+ — this is a new image that resembles the old one.

If your generative upscale looks wrong, it is almost always denoise set too high. That is the one setting to learn.

Many of these take a prompt, and it matters. Describe what is in the image accurately and the synthesised detail will be appropriate; leave it empty and the model guesses; describe it wrongly and it will confidently add the wrong thing.

Tiling, and the seam problem

Upscaling to 4K+ will not fit in VRAM as one operation, so the image gets split into tiles, upscaled separately, and reassembled.

Two consequences:

Seams. Tiles processed independently disagree at their edges. The fix is overlap — processing tiles with a margin and blending across it. In ComfyUI’s tiled upscale nodes this is a tile_overlap or similar parameter; 64 to 128 pixels is a reasonable range. Too little gives you visible grid lines; too much wastes compute.

Tile-level context loss. A generative upscaler working on a 512px tile cannot see the whole image, so it does not know that this patch of brown is part of a dog. With a prompt, it applies the prompt’s description to every tile — which is why you occasionally get a small extra eye in the corner of a generative upscale. Lower denoise reduces this; so does a more generic prompt.

Video: the extra problem is time

Upscaling video frame-by-frame with an image upscaler produces temporal flicker — each frame’s invented detail differs, so surfaces boil. It is unwatchable for anything longer than a few seconds.

Options, in order of practicality:

Use a video-native upscaler. Models built for video enforce consistency across frames. These are the right answer and they are heavier.

Use a low denoise and a deterministic seed. A fixed seed plus minimal generation keeps frames closer together. It reduces flicker rather than eliminating it.

Upscale non-generatively. Lanczos or Real-ESRGAN at low strength flickers much less than a diffusion upscaler, because it invents much less.

Or deflicker afterwards with a temporal smoothing pass, which costs sharpness.

When not to upscale

When the image is evidence. Archival, documentary, journalistic, forensic, scientific. A generative upscale of a document, a face or a specimen is a fabrication that looks like a photograph. If you do it for presentation, label it.

When you are about to print. Check what your printer actually needs. 300dpi at final size is the usual rule, and a surprising amount of work that people upscale 4x never needed it.

When the source is compressed. Upscaling a heavily JPEG-ed image amplifies the block artifacts into confident structure. Denoise or de-block first, then upscale.

And when you could regenerate instead. If the image came from a generative model, re-running it at a higher base resolution is almost always better than upscaling the small one. The model had the latent; use it.