Audio machine learning and creative sensing
Stem separation
How models pull drums, bass and vocals out of a mix, and where they struggle
Early draft. We're sharing this guide for review today. More visuals, sounds and playable examples are on the way; boxes marked "Lab" describe the interactive parts we're building.
Lab: Take the song apart · Inspect a prepared result
- Starts with: a short, cleared song playing in full, with four muted lanes underneath: Vocals, Drums, Bass, Other.
- Reader controls: solo or mute each lane while the song plays.
- Ask before playing: "If you solo the vocal lane, what do you think you'll hear besides the voice?"
- Fallback: four short clips and a spectrogram image of the mix and each stem.
- Owner: modules/stemsep (pre-separated result for instant playback, labelled "separated in advance with StemSep").
Play the song, then solo Vocals. The singer is suddenly alone, as if the band walked out of the room. Solo Drums: there's the beat, with no voice over it. Mute the vocal and you've got an instrumental.
It feels like magic, because this song was mixed down into a single stereo file. There are no separate tracks hidden inside it. So where did the stems come from?
Try it: separate it yourself
Lab: Run StemSep · Inspect a prepared result
- Starts with: StemSep loaded with the same song; nothing separated yet.
- Reader controls: Separate; then the four stem lanes with solo, mute and level.
- Shows: progress while it runs; afterwards, the four stems as waveforms.
- Note shown to the reader: "The model downloads once (about 79 MB) and runs in your browser. Your audio stays on your device."
- Fallback: the pre-separated result from the first lab.
- Owner: modules/stemsep (Spleeter ONNX, fp16 by default; fp32 optional).
Press Separate. The first time, your browser downloads the model; after that it's cached. Then it works through the song and four stems appear.
Now listen closely, the way an engineer would. Solo the vocal and listen in the gaps between sung lines. Can you hear a ghost of the hi-hat? A wash of reverb from the guitar? Solo the bass and listen to the very start of each note. Is the attack as sharp as you'd expect?
What just happened
The model never found the original tracks, because they no longer exist. When a song is mixed, every instrument is added together into one signal, and there's no way to un-add them exactly. What the model does is estimate each source.
Here's the idea, in plain terms:
- The audio is turned into a spectrogram, a picture of which frequencies are present at each moment.
- The model has been trained on thousands of songs where the separate tracks were available. From those, it learned what vocals, drums and bass tend to look like in that picture.
- For each stem, it paints a mask over the spectrogram: for every little time-and-frequency tile, how much of it probably belongs to the vocal, how much to the drums, and so on.
- Each mask is applied to the mix and turned back into audio. Those are your stems.
What's a mask? (optional): imagine the spectrogram as a grid of tiles. The vocal mask gives each tile a number from 0 to 1. A tile where only the voice is playing gets close to 1. A tile where only the hi-hat is playing gets close to 0. A tile where the voice and the guitar overlap gets something in between, and that's where the trouble starts.
That's why we call them estimates. Each stem is the model's best guess, shaped by what it saw during training.
Where it breaks
Lab: Estimate versus original · Listen and compare
- Starts with: the same song, with the original studio stems available alongside StemSep's estimates.
- Reader controls: for each instrument, switch between Original and Estimate (level-matched).
- Ask: "Which stem do you think the model got closest? Which one furthest?"
- Fallback: paired clips and a short table of what to listen for.
- Owner: modules/stemsep estimates; originals from the cleared multitrack (to be sourced).
Now you can hear exactly what the model got wrong, because you have the real thing to compare against. Listen for these:
- Leakage (bleed). Bits of one instrument in another's stem. Hi-hats and cymbals often leak into the vocal because they share high frequencies with sibilant "s" sounds.
- Smeared attacks. Drum hits and plucked bass notes can lose some of their sharp front edge, because a spectrogram trades time detail for frequency detail.
- Watery or "underwater" artifacts. When the mask flickers between keeping and dropping the same tile, the result can warble.
- Reverb that doesn't know where it belongs. A reverb tail is a blend of everything; the model has to split it somehow, and often splits it oddly.
- "Other" is a catch-all. Guitars, keys, synths, strings and backing vocals all land in one stem. The model was trained on four categories, so it can't give you a separate piano.
One more experiment: unmute all four estimates together and switch between them and the original mix. It's close, but listen for what's different.
Our take: separated stems are excellent for practice, sampling, DJ edits, remixes and studying how a mix was built. For a release where one stem will be heard exposed and alone, test it carefully, and expect to do some cleanup.
The model behind it
StemSep in Maru runs Spleeter, an open-source separation model released by Deezer in 2019, converted to run in the browser. You can choose two versions: fp16, a smaller download (about 79 MB), and fp32, about twice the size (about 157 MB). They're the same model stored at two numeric precisions, not two different methods. Try both on the same passage and decide whether you hear a difference.
Newer models work differently, for example on the waveform directly or on both the waveform and spectrogram, and report higher scores on the standard public benchmark (MUSDB18) than Spleeter. Benchmark scores are averages over a test set of songs; they won't tell you how a model handles your song. Your ears and a comparison with a known original will.
Same idea, different place: the spectrogram the model works on is the same picture you learned to read in Seeing sound, and its time-versus-frequency trade-off is why the attacks smear.
Make it yours
- Separate a different song of your choice (up to three minutes). This time you don't have the originals, so you're the judge.
- Solo each stem and write down one place where it breaks: a leak, a smeared hit, a watery patch.
- Build a short edit: mute the vocal for a section, bring up the bass, add a delay throw to one vocal phrase.
- Export the stems you used as WAV files, or open them in the studio.
Make sure you have the rights to use the song the way you plan to. Separating a track for practice is different from releasing a remix.
[Open these stems in the studio] (planned link: loads the separated stems as tracks in a studio project)
Quick check
- A friend says "the splitter recovered the original vocal track". What's
more accurate?
Answer
It produced an estimate of the vocal, based on what it learned from training data. The original track isn't in the mixed file. - You hear faint cymbals in the separated vocal. What's that called, and why
is it common?
Answer
Leakage or bleed. Cymbals share high frequencies with vocal sibilance, so the mask can't cleanly tell them apart. - You need a clean, isolated piano for a sample. Will this model give you
one?
Answer
Not on its own: it only separates vocals, drums, bass and "other", and piano lands in "other" with everything else.
Next steps
- Go deeper: How models see audio. Spectrograms, latents and tokens, and why the picture a model uses decides what it can do.
- Use it: Making room with EQ. Clean up leakage and make your remix stems sit together.
- Build it: Adding a model to your app. How StemSep downloads, caches and runs a model in the browser, and how to do the same.
Go deeper
Inside StemSep. It accepts mono or stereo clips up to three minutes or 128 MiB. Model downloads are checked against a recorded checksum and size, then cached in the browser; Clear model cache removes them. The local methods keep uploaded audio on the device. Results can be exported as individual WAV files or a ZIP with a record of the run.
Sources and examples. The demo song must come with its original stems under a licence that allows this use; to be sourced and recorded before publishing. Spleeter: Deezer Research, 2019 (link and licence to add). MUSDB18 benchmark context to be linked to a primary source before publishing.