Troubleshooting
Fixing the Most Common Stem Separation Artifacts
Separation artifacts are not random. Each one has a specific cause in how the model assigns energy, and each has a repair that works — plus several popular repairs that make things measurably worse.
Every separated stem contains some amount of estimation error, and that error shows up in a small number of recognisable ways. Once you can name what you are hearing, the fix is usually straightforward. This guide walks through the artifacts in rough order of how often they cause trouble, with the causes and the repairs that hold up.
One rule applies throughout, so it is worth stating first: do not re-run a separated stem through the separator to clean it up. The model expects a full mix as input. Feeding it a stem that already has errors produces a result that estimates on top of those errors, and the artifacts compound. When a stem is unusable, go back to the original file and change the mode.
Vocal ghosting in the instrumental
What you hear: a faint, breathy shadow of the original singer under the band, most obvious on held notes and in sparse verses.
Why: the mask for those time-frequency cells landed in the middle rather than near 0 or 1, so the energy was split between stems. Heavily compressed masters make this much worse, because loudness processing glues the vocal to the rest of the mix and reduces the contrast the model relies on.
Fixes, in order of what to try first:
- 01Switch to two-stem mode if you were using five. A dedicated voice/not-voice split usually leaves less residue than summing four estimated stems.
- 02Find a better source. A CD rip or purchased WAV routinely fixes ghosting that no amount of processing will.
- 03Apply a dynamic EQ with a wide band centred around 2–3 kHz, threshold set so it only engages when the ghost appears. This preserves the instrumental's tone between vocal phrases, which a static EQ cut cannot.
- 04If the ghost is centred and the mix is wide, attenuate the mid channel in the 300 Hz–5 kHz range with a mid/side tool. Watch the kick and snare, which also live in the centre.
- 05Mask it. In practice, a live singer or a new vocal over the top hides residue far more effectively than processing does, and without collateral damage.
Advertisement
Watery, phasey, metallic vocals
What you hear: the isolated vocal swirls, as though it is being played through a slow flanger. Often described as underwater or robotic.
Why: the mask values fluctuate between adjacent analysis frames, so individual harmonics flicker in and out from one moment to the next. Your ear integrates that flickering as a moving comb filter. It is worst on quiet passages and on vocals with heavy reverb, because the reverb tail is spread thinly across many uncertain cells.
- Do not try to EQ this out. The problem is temporal, not spectral, and EQ cannot address it.
- Roll off below about 100 Hz and above about 12 kHz. Most of the audible swirl lives in the extremes where the vocal has little genuine content, so filtering removes artifact without removing voice.
- Add reverb. This sounds like a cheat and it is the single most effective repair: a short plate or room reverb smears the flickering into a continuous tail and the swirl stops reading as an artifact. If the vocal is going into a new production this is free, because you were going to add reverb anyway.
- For dialogue rather than singing, a light broadband noise reduction at 3–6 dB can smooth the flicker without the pumping you get from aggressive settings.
- If the vocal must be pristine, check whether the song has a released acapella or a stem pack. No amount of processing beats the real thing.
Hi-hats and cymbals in the vocal stem
What you hear: a rhythmic tick or shimmer riding along with the isolated vocal.
Why: this one is a genuine ambiguity rather than a failure. A closed hi-hat and a sung or spoken sibilant are both short bursts of broadband high-frequency noise with sharp onsets. On a spectrogram they can look nearly identical. The model resolves the ambiguity using surrounding context, and when the context is weak it guesses.
- A high shelf cut above 8–10 kHz reduces the tick noticeably and costs you some sibilance. Usually a good trade for practice and transcription work.
- A transient shaper set to soften attacks reduces the perceived tick while leaving the vocal body intact, since sung notes have gentler onsets than hi-hats do.
- If you are working in a DAW, the cymbal bleed is periodic. Automating a narrow cut on the grid is tedious but surgical, and worth it for a small number of bars.
- Where the goal is an instrumental rather than a vocal, this artifact is harmless — the hats are still present in the instrumental too.
Bass that has lost its punch
What you hear: the bass stem has the right notes but feels soft, and the low end sits behind the beat rather than driving it.
Why: the perceived punch of a bass note comes largely from its attack, which contains mid and even high-frequency energy — the click of a pick, the thump of a string against a fret. The model tends to assign the sustained fundamental to bass and the attack transient elsewhere, most often to drums or other. You are left with the note but not its beginning.
- Layer the bass stem with the drum stem when monitoring. In most cases the missing attack is sitting in the drum file, and hearing them together restores the groove.
- A transient designer with attack boosted 2–4 dB partially rebuilds the click. Do not push it far or you will amplify artifacts along with the transient.
- For production use, trigger a sample or a synth sub from the bass stem rather than trying to repair it. The stem is an excellent pitch and timing reference even when its tone is compromised.
- Check whether the song's bass is synthesised. Sub-heavy electronic bass often separates poorly because it has almost no harmonic structure for the model to latch onto; three-stem mode, which keeps bass inside the instrumental, may serve you better.
Reverb and delay tails in the wrong place
What you hear: the dry vocal is in the vocal stem, but its reverb tail is faintly audible in the instrumental — or vice versa, so the isolated vocal sounds oddly dry and clipped at the end of each phrase.
Why: reverb is a diffuse, decaying copy of a source spread across time and frequency, and by the time it has decayed it no longer resembles the thing that caused it. The model routinely assigns the direct sound and its tail to different stems.
- For a too-dry isolated vocal, add your own reverb. Matching the original is a guessing game; a plate with a decay in the 1.2–1.8 s range is a reasonable starting point for most pop material.
- For tails left in the instrumental, treat them like ghosting and use dynamic EQ. They are usually quiet enough to ignore once anything else is playing.
- If a song is drenched in reverb — a lot of shoegaze, ambient, and live recordings — accept that clean separation is not available. The information is genuinely entangled.
An instrument in the wrong stem entirely
What you hear: the electric guitar is in the piano stem. The strings are in other. The organ is split across three files.
Why: the five-stem model knows four named instrument classes plus a remainder. Anything it was not trained to recognise gets sorted into whichever class it most resembles spectrally. This is expected behaviour, not a bug.
- Sum the stems you do not need back together and work with the grouping the model can actually deliver. Piano plus other is often the coherent harmonic bed you were looking for.
- Drop to three stems, which avoids the distinction entirely by keeping all pitched material in one file.
- For a specific instrument you need cleanly, check whether it occupies a distinct frequency range you can isolate with EQ after separation. A high organ line or a low cello part is sometimes easier to extract by filtering than by class.
Clicks, dropouts and gaps
What you hear: brief silences or clicks at seemingly random points, sometimes at consistent intervals.
Why: this is usually not the model. It is more often the source file — a corrupted download, a variable-bitrate MP3 with a damaged frame, or a file that was cut from a stream mid-packet. Because the separator reconstructs audio from the input's own phase information, a defect in the source propagates into every stem at the same position.
- Play the original file end to end first. If the click is in the source, it will be in all your stems.
- Re-encode the source to WAV before uploading. This often repairs container-level problems in MP3 and M4A files.
- If the gap appears in only one stem and not the source, it is a mask dropout. Crossfade in a fragment from the surrounding audio, or accept it if the gap is under a few milliseconds.
The repairs that do not work
Some approaches circulate widely and consistently make things worse. Worth knowing so you do not waste an afternoon:
| Popular idea | Why it backfires |
|---|---|
| Separating the same file twice for a cleaner result | The second pass treats artifacts as signal. Errors compound rather than cancel. |
| Aggressive broadband noise reduction on a swirly vocal | The artifact is not stationary noise. Strong settings produce pumping and chew the consonants. |
| Boosting the top end to bring back air | If the source was a lossy MP3 there is nothing above roughly 16 kHz to boost except encoder noise. |
| Phase-cancelling the instrumental against the original to get a cleaner vocal | Only works when the instrumental is a true bit-accurate complement. A separated instrumental is not, so you get comb filtering. |
| Normalising each stem to full scale | Destroys the relative balance that made the stems useful together, and pushes artifacts up along with signal. |
When to stop
The most valuable judgement in this work is knowing when a stem is good enough for its purpose. A vocal with mild swirl is completely fine as a transcription reference or a practice guide. An instrumental with a faint ghost is fine for a karaoke night. Neither is fine for commercial release, and no chain of plugins will get them there. If the intended use demands genuinely clean audio, the answer is a licensed stem pack or a re-recording, not more processing.
Try it on your own track
VocaSplitter splits a song into vocals, drums, bass, piano and other stems in your browser. No account, no watermark, no cost.
Open the separator