Stem separation went from impossible to boring in about five years. Pulling a clean vocal out of a stereo mix used to be a party trick that didn’t really work; now it is a checkbox in most DAWs and a paid tier in most online tools.
Demucs is the open-source model underneath a great deal of that, it runs on your own machine, and it costs nothing. If you are paying a subscription to separate stems, this is very likely the thing you are paying a wrapper around.
musictechtuition walks through Demucs GUI and compares its output against SpectraLayers and RipX.
What it does
Demucs takes a stereo mix and outputs four files: drums, bass, vocals, other. Some model variants add a guitar and piano split, at some cost to the other four.
It is a waveform-domain model — it works on the audio samples rather than only on a spectrogram — which is the architectural reason its output tends to sound less “phasey” than older spectrogram-masking approaches.
The two ways to run it
Demucs GUI (start here)
If you don’t want to touch a terminal, Demucs GUI is a packaged application that bundles the models. Download it, point it at a file, pick a model, press go. This is what the video above covers, and for most people it is the correct answer — the command-line version offers control you probably don’t need yet.
Command line
If you have Python 3.9 or newer:
python3 -m pip install -U demucs
Then, simplest possible use:
demucs mytrack.wav
Output lands in separated/htdemucs/mytrack/ as four WAV files.
Useful flags:
# Extract only the vocals, and one "no_vocals" backing track
demucs --two-stems=vocals mytrack.wav
# Higher quality, ~4x slower (bag of four models)
demucs -n htdemucs_ft mytrack.wav
# Six stems: adds guitar and piano
demucs -n htdemucs_6s mytrack.wav
# MP3 output instead of WAV
demucs --mp3 --mp3-bitrate 320 mytrack.wav
# Force CPU if the GPU path misbehaves
demucs -d cpu mytrack.wav
Picking the model
htdemucs— the default. Fast, good. Use it unless you have a reason not to.htdemucs_ft— fine-tuned, four passes, roughly four times slower and audibly better. Worth it on anything you’ll actually release.htdemucs_6s— six stems. The guitar and piano splits are experimental and the added separation slightly degrades the other four. Try it, don’t assume it.mdx_extra— trained differently; sometimes wins on material the others struggle with. Worth a comparison pass on a difficult track.
Speed
On a CUDA GPU a three-minute track takes seconds to tens of seconds. On CPU, expect minutes — and on Apple Silicon, -d mps sometimes works and sometimes produces artefacts depending on your PyTorch version; if the output sounds wrong, fall back to -d cpu before blaming the model.
If you run out of GPU memory, reduce the segment size:
demucs --segment 8 mytrack.wav
What to listen for before you build on the output
This is the part tutorials skip, and it is the part that determines whether your finished work sounds amateurish.
Cymbals bleed into “other.” Consistently. High-frequency transient content is the hardest thing to assign, and hi-hats and crash decays end up smeared across the drums and other stems. If you are rebuilding a mix from stems, this shows up as a strange top end.
Reverb tails go with the wrong stem. A vocal’s reverb was printed into the mix; the model has to decide whether that wash belongs to the vocal or the room. It often splits the difference, which is why isolated vocals frequently sound drier than they should and the “other” stem has a ghost of the vocal’s ambience.
Dense mixes degrade more than sparse ones. A four-piece rock recording separates well. A wall-of-sound production with layered synths and heavy bus compression separates much less well, because the model is trying to undo a deliberate blend.
Solo the stems in isolation, at volume, before committing. Artefacts that are inaudible inside a full mix become obvious the moment a stem is exposed — and if you are sampling a stem into something new, exposed is exactly where it will end up.
What it’s genuinely good for
- Remixing and sampling — the original use case, and where imperfect separation matters least because you are burying the result in new material
- Practice and transcription — kill the bass to play along, isolate a guitar part to work out a line
- Karaoke and live backing tracks —
--two-stems=vocalsis a one-command solution - Rescuing your own archives — old bounces where you’ve lost the session files
- Analysis — feeding separated stems into transcription or beat-tracking tools works markedly better than feeding a full mix
The part worth saying plainly
Separating a commercial track is technically trivial and legally not your call. The model doesn’t know or care what you feed it; releasing something built from a copyrighted recording is a rights question that has nothing to do with which tool you used. Practice, transcription, and your own material are uncomplicated. Anything you publish is not.