The fastest-rising repository by likes on our Hugging Face tracker this morning is not a foundation model. It is akatz-ai/MiniMax-H3-Character-Swap-LoRA — up 44% in likes in a single day, 99 to 143, on 5,890 downloads. Published 25 September, updated 27 September.
It does one thing: takes a video and a reference image, and replaces the character in the video with the one in the image.
Patrick Sheedy walks through the ComfyUI workflow end to end.
What it is technically
A LoRA adapter on Comfy-Org/MiniMax-H3, distributed as a single diffusion file — which is the practical detail that explains the adoption curve. One file, into your LoRAs folder, loaded by an ordinary LoRA node. No custom nodes to install, no repository to clone, no environment to rebuild.
The tags describe the shape precisely: video-to-video, ref2va (reference-to-video-and-audio), character-swap, comfyui, ai-toolkit. It was trained on akatz-ai/H3-Character-Swap-v1, a published dataset — which is more disclosure than most LoRAs of this kind offer, and worth noting in a category where training data is usually silent.
Why the source video matters more than the reference image
The instinct is to spend your effort on the reference image. That is backwards. The source video is doing almost all of the work, and it is where the failures come from.
A swap of this kind preserves motion, framing and timing from the source and re-renders identity on top. So the source dictates:
- What poses the model must handle. A face turning past profile, a hand crossing the face, a fast whip-pan — these are where identity breaks, and no reference image fixes them.
- The lighting the new face has to sit in. The model is matching the source’s light. A source with hard directional light and strong colour cast is a harder problem than flat, even lighting, and it is where the seam shows.
- How long the shot is. Longer generations drift. Identity consistency degrades over time in essentially every video model, and it is the reason the ecosystem around MiniMax H3 has produced a separate body of work on long-video latent context.
The corollary for anyone using this in real work: shoot for the swap. Even, diffuse lighting, moderate motion, faces kept mostly frontal, shots kept short and cut together afterwards. The constraints resemble the ones that governed practical compositing before anyone used a diffusion model, which is not a coincidence.
Where it fits in a real pipeline
Not as a finishing tool. As a previsualisation and iteration tool, it is genuinely useful:
- Testing a character design against real performance before committing to a full 3D or 2D pipeline
- Trying five different castings of an animatic without reshooting
- Stand-in performance capture where the performer is not the character
- Style and identity experiments where the motion is already right and only the appearance is in question
That last case is the honest sweet spot. Getting convincing motion is the hard part of generated video, and it remains hard. Taking motion you already have — filmed, or animated, or from a previous generation — and changing only who is performing it sidesteps the part the models are worst at.
The part that should be said plainly
Replacing the person in a video with a different person, from a single photograph, in a free download that runs on consumer hardware, is a deepfake tool. Calling it a character swap is accurate and also not the whole description.
The technology is out and has been for a while; a LoRA is not the thing that changed that. But if you are publishing work made with it, label it, and if the face you are putting into a video belongs to a real person, get their permission — a requirement that is legally enforceable in a growing number of jurisdictions and ethically obvious in all of them. The interesting creative uses here are overwhelmingly ones where the character is invented, which is also the uncomplicated case.