AI & Creative Tools

ControlNet: Getting Started

Prompting tells a model what to make. ControlNet tells it where to put things — pose, depth, edges, layout — so you stop rerolling seeds and start directing.

The gap between “I like this image” and “I need this image” is composition. Prompting is bad at composition — you can describe a pose for twenty minutes and get twenty different poses. ControlNet closes that gap by letting you hand the model a picture of the structure you want.

1. What ControlNet actually is

A diffusion model normally has one steering input: your text. ControlNet adds a second — a spatial one.

It works by running a parallel copy of part of the model’s network, conditioned on a control image, and injecting that conditioning at each denoising step. The practical effect: the text decides what things are, the control image decides where they go.

The critical thing beginners get wrong: the control image is not your reference photo. It’s a processed abstraction of it — a stick-figure skeleton, a greyscale depth map, a white-on-black edge drawing. The step that makes it is called a preprocessor, and choosing the right one is the skill.

2. The five worth learning first

PreprocessorWhat it extractsUse it when
OpenPoseA skeleton: joints and limbs, optionally hands and faceYou care about the pose and nothing else about the reference
DepthA greyscale map, near = light, far = darkYou want the 3D arrangement and camera feel preserved
CannyHard black-and-white edgesYou want tight fidelity to shapes and outlines
LineartCleaner, more drawing-like lines than CannyColouring line art, or illustration-style work
ScribbleLoose, thick strokesYou’re drawing the layout yourself, roughly

The ordering in that table is roughly loosest to tightest control. OpenPose throws away everything except the figure’s posture, so the model is free to invent clothing, body type, setting. Canny keeps nearly every contour, so the model is mostly colouring in. Neither is better; they answer different questions.

Pick the loosest one that captures what you actually care about. This is the single highest-leverage habit. If you only need the pose, using Canny hands the model the reference’s hair, clothes and background too — and then you fight to get rid of them.

3. The minimum workflow

In ComfyUI, the graph is:

Load Image  →  ControlNet Preprocessor  →  Apply ControlNet  →  KSampler
                                              ↑
                    Load ControlNet Model ─────┘
                    CLIP Text Encode (pos/neg) ┘

You need a ControlNet model file matching your preprocessor and your base model — a depth ControlNet trained for SDXL will not work with an SD1.5 checkpoint, and vice versa. Mismatched pairs are the most common reason “ControlNet does nothing.”

Then: load your reference, run the preprocessor, look at the preprocessor output before you generate. If the skeleton has the arms wrong, or the depth map is flat mush, nothing downstream can save it. Half of ControlNet debugging is just looking at the intermediate image.

4. The three parameters that matter

Strength (0.0–2.0, default 1.0) — how hard the control pulls.

  • 0.4–0.7: suggestion. The model takes your structure as guidance and deviates.
  • 0.8–1.0: the normal working range.
  • 1.2+: rigid adherence, and the point where images start looking stiff and traced.

Start percent / End percent — when during the denoising the control applies, as a fraction of total steps.

This is the underused one. Composition is decided in the early steps; detail happens in the late ones. So:

  • start 0.0, end 0.5 — control the layout, then let the model finish freely. Usually the best-looking results.
  • start 0.0, end 1.0 — control throughout. Faithful, often flatter.
  • start 0.2, end 0.8 — let the model establish its own broad shapes first, then constrain.

If your ControlNet output looks correct but lifeless, lower the end percent before you lower the strength. Releasing the control for the final third of the steps is usually what restores the texture and light the model is good at.

Resolution — the preprocessor’s output resolution. Match it roughly to your generation size; a 512px depth map upscaled to a 1536px generation is why your edges are soft.

5. Stacking

You can chain multiple ControlNets — Apply ControlNet nodes feed into each other.

The combination that earns its keep: OpenPose + Depth. Pose fixes the figure, depth fixes the space it’s standing in, and neither dictates appearance. Run both around 0.6–0.8 strength; stacked controls at full strength fight each other and produce mangled output.

Avoid stacking two tight controls (Canny + Lineart, say). They’re describing the same information twice and the model has no freedom left.

6. Why your results look pasted-on

Four usual causes, in order of frequency:

  1. Strength too high. Drop to 0.7 and see what happens.
  2. Control applied for the full duration. Set end percent to 0.5–0.6.
  3. Wrong preprocessor. You used Canny when you meant OpenPose, and inherited the reference’s lighting and clothing as edges.
  4. Prompt fighting the control. If the skeleton is sitting and the prompt says “standing,” you get a sitting figure with standing-figure anatomy. The control wins on geometry; the prompt wins on everything else. They have to agree.

7. Where to take it

Once single-image ControlNet is comfortable, the two obvious directions are posing a 3D figure yourself — several tools let you arrange a mannequin and export an OpenPose skeleton directly, which beats hunting for a reference photo — and video, where running a consistent depth or pose control across frames is the foundation of most controllable AI animation workflows.