Guides
2, 3 or 5 Stems: Choosing the Right Separation Mode
Every additional stem is another chance for the model to guess wrong. Pick the fewest stems that answer your actual question, and your audio will sound better.
The separation mode selector is the one choice you make before processing, and most people default to the highest number on the assumption that more stems means more capability. It does mean more capability — but not more quality, and the trade is steeper than it looks.
What each mode produces
| Mode | Files you get | Best for |
|---|---|---|
| 2 stems | Vocals, Instrumental | Karaoke, acapellas, removing a voiceover, anything where you only care about voice versus everything else |
| 3 stems | Vocals, Drums, Instrumental | Drum practice and transcription, DJ edits where you need to loop or replace the beat |
| 5 stems | Vocals, Drums, Bass, Piano, Other | Remixing, rebalancing a mix, learning a specific instrument part, sample sourcing |
Why fewer stems sounds better
The model works by deciding, for every tiny slice of time and frequency in the song, which stem that energy belongs to. In two-stem mode it answers one question: voice or not voice. In five-stem mode it answers a five-way question about the same slice, and the boundaries between some of those categories are genuinely blurry — a piano and a synth pad occupy overlapping harmonic space, and a bass guitar's attack contains mid-range energy that looks a lot like a muted guitar.
Two consequences follow. First, each stem in a five-way split carries more error than either stem in a two-way split. Second, if you sum all five stems back together to make an instrumental, you inherit every one of those errors at once, and any energy the model assigned confidently to nothing at all is simply missing. The two-stem instrumental is almost always fuller and more coherent than a rebuilt one.
Advertisement
When two stems is right
Reach for two stems whenever the voice is the axis you care about. Karaoke and backing tracks, obviously. Also: extracting an acapella for a mashup, pulling a narrator off a documentary bed so you can re-dub it, isolating a rap verse to study its cadence, or making an instrumental version of your own demo when you have lost the session files. The two-stem model is also the most robust on difficult material — if a dense modern pop track is falling apart in five-stem mode, two stems will often still give you something usable.
When three stems is the sweet spot
Three stems is underrated. It adds the single most useful extra class — drums — while keeping everything harmonic in one coherent bucket, so you avoid the piano-versus-other confusion entirely.
It is the right choice for drummers learning a groove, for anyone transcribing a beat, and especially for DJs and producers building edits: you can loop the drum stem to extend an intro, drop the drums out to create a breakdown, or replace the kit entirely while keeping the harmonic content whole. Percussion also happens to be the class models separate most reliably, because a transient burst of broadband noise is acoustically very distinct from sustained pitched material.
When five stems earns its cost
Five stems is the mode to use when you need to see inside the arrangement rather than around the voice. Real cases where it is clearly correct:
- Working out a bass line by ear, where hearing it without the kick drum is the whole point.
- Rebalancing a mix you cannot re-open — pulling a too-loud piano down 3 dB and pushing the drums up.
- Sourcing a specific sample: a piano chord, a bass note, a percussion loop.
- Studying an arrangement to understand how the parts interlock.
- Reharmonisation work, where you need the pitched material separated from the rhythm section.
Expect the piano and other stems to be the weakest of the five. If the song has no piano, the piano stem will not be empty — it will contain whatever mid-range harmonic material most resembled a piano to the model, often a rhythm guitar or a pad. That is not a malfunction; it is the remainder problem described above.
How the source material shifts the answer
The best mode depends partly on what you are separating.
- Sparse acoustic material — voice and one or two instruments — separates well in every mode. Use five if you want the detail; the penalty is small.
- Dense, loud, modern productions get worse quickly as you add classes. Start at two, only go higher if two was clean.
- Electronic music with synthesised bass often confuses the bass and other classes, since a sub-bass synth and a bass guitar have quite different spectral shapes. Three stems is frequently more useful than five here.
- Anything built on instruments outside the model's four learned classes — brass sections, strings, steel pan, talking drum, mandolin — will pile into other. Five stems mostly tells you what is not those instruments.
- Speech content such as podcasts and interviews should use two stems. The music classes have nothing to do.
A practical workflow
If you are unsure and the song matters, do this: run two stems first and listen to the instrumental. That takes a couple of minutes and immediately tells you how cooperative the material is. If the two-stem result is clean, a five-stem split of the same file will also be reasonable and you can go get the detail. If the two-stem instrumental already has audible ghosting, five stems will be considerably worse, and you have saved yourself the disappointment.
One thing not to do: separate a stem and then separate that stem again to break it down further. Each pass estimates on top of the previous pass's errors, and the artifacts compound audibly. Always go back to the original mix and choose a different mode.
Try it on your own track
VocaSplitter splits a song into vocals, drums, bass, piano and other stems in your browser. No account, no watermark, no cost.
Open the separator