Frequently asked questions
The questions we actually get asked, grouped by topic. If something is missing, the guides go into more depth, or you can email us.
Using the tool
Is VocaSplitter really free, and what is the catch?
It is free with no upload limit, no account and no watermark. The costs are covered by display advertising on the written guides. There is no paid tier holding back better quality — you get the same model output either way.
Do I need to create an account?
No. There is no sign-up, no email address required and no login. Upload a file and you get stems back in the same session.
How long does a separation take?
Usually one to three minutes for a three-minute song, depending on how busy the queue is and which mode you chose. Five-stem separations take longer than two-stem ones. Longer files scale roughly linearly.
Is there a file size or length limit?
There is no hard cap we enforce, but very long files (over about fifteen minutes) are more likely to time out because they hold the processing slot for a long time. If you need to separate a full DJ set or a podcast episode, split it into shorter sections first and separate each one.
Which audio formats can I upload?
MP3, WAV, M4A, AAC, OGG, FLAC and AIFF all work. If a file is rejected, re-exporting it as a 44.1 kHz WAV fixes most problems — rejections are usually caused by a damaged container rather than the format itself.
What format do the stems come back in?
WAV, at the sample rate of the file you uploaded. WAV is lossless, so nothing is degraded a second time by an encoder. Convert to MP3 yourself afterwards if you need smaller files.
Can I use it on a phone?
Yes. The interface works on mobile browsers and the processing happens on our server, so your device only needs to upload the file and play the results. Downloading multiple stems is generally easier on a desktop.
Can I preview stems before downloading?
Yes — every stem has a play control on the results screen. Preview first and download only what you need, or take everything as a single ZIP.
Advertisement
Quality and results
Which separation mode should I choose?
Use the fewest stems that answer your question. Two stems for a karaoke instrumental or an acapella, three if you also need the drums out, five when you genuinely need the individual parts. More stems is not better quality — each extra class is another decision the model can get wrong.
Why can I still hear the singer faintly in the instrumental?
That is vocal ghosting, and it happens when the model was not confident about some of the audio, so it split that energy between stems. Loud, heavily-compressed masters make it worse because the vocal is glued to everything else. A better source file helps more than any processing, and two-stem mode usually leaves less residue than rebuilding an instrumental from five stems.
Why does the isolated vocal sound watery or metallic?
The estimate fluctuates slightly between adjacent moments in time, so individual harmonics flicker and your ear reads that as a swirl. Adding a short reverb is the single most effective fix because it smears the flicker into a continuous tail. EQ will not help — the problem is in time, not frequency.
There are hi-hats in my vocal stem. Is that a bug?
No, it is a real ambiguity. A closed hi-hat and a sung 's' are both short bursts of broadband high-frequency noise and can look nearly identical on a spectrogram. A high shelf cut above 8 kHz reduces the tick at the cost of some sibilance.
Which songs separate best?
Sparse arrangements with one clear lead vocal — singer-songwriter material, soul, gospel, most pre-2000 rock. Dense modern pop with layered synths and heavy limiting is hardest, as are heavily autotuned or hard-doubled vocals, live recordings with crowd bleed, and anything drenched in reverb.
Does the quality of my source file matter?
More than any setting. A 128 kbps MP3 has already lost most content above roughly 16 kHz, and the model will faithfully separate a song with no top end. Use WAV or FLAC where you can, and at minimum a 256 kbps MP3.
Can I run a stem through the tool again to clean it up?
Please do not — it reliably makes things worse. The model expects a full mix, so a second pass treats the first pass's artifacts as signal and compounds them. Go back to the original file and try a different mode instead.
Why is there something in the piano stem when the song has no piano?
The five-stem model knows four named instrument classes plus a remainder called 'other'. Anything it was not trained to recognise gets sorted into whichever class it most resembles, so a rhythm guitar or a synth pad often lands in the piano stem. Three-stem mode avoids the distinction entirely.
Are the stems in sync with each other?
Yes, to the sample. Every stem is cut from the same source file and has the same length, including any leading silence. Drop them all at the session start in your DAW and they line up exactly — do not trim one before importing.
Is the quality good enough to release commercially?
Usually not for an isolated vocal — you can hear the process on headphones if you listen for it. For karaoke, practice, transcription, DJ edits, rehearsal material and sample sourcing it is comfortably good enough. If you need genuinely clean audio for release, a licensed stem pack or a re-recording is the honest answer.
Privacy and rights
What happens to the file I upload?
It is sent over HTTPS to the processing server, separated, and then deleted. We do not keep a library of uploads, do not train models on your audio, and do not pass files to third parties. Download links are short-lived, so save your stems during the session.
Do you store the stems you generate?
No. Generated stems are removed along with the source file after processing. If you come back tomorrow the files will be gone and you will need to separate again.
Am I allowed to separate a commercial song?
Separating a recording creates a derivative work. Doing that for private study or practice is generally uncontroversial. Publishing the result, performing it in a venue, or monetising it typically needs permission from whoever controls both the recording and the composition. Our guide on separation and copyright covers where the lines fall.
Can I use the stems in a YouTube video or a released track?
Only with the appropriate rights. Content matching systems still recognise a separated instrumental, because its fingerprint remains close to the original recording — removing the vocal does not prevent a claim. For released work, use licensed stems, clear the sample, or re-record the parts yourself.
Do you use cookies or show personalised ads?
The site uses Google AdSense, which may set cookies to measure and serve ads. The tool itself does not require cookies to function. Details, including how to opt out of personalised advertising, are in the privacy policy.
I am a rights holder and want something addressed.
Email vocasplitter@gmail.com and we will respond. Note that uploads and generated stems are deleted after processing, so nothing is hosted or indexed here on an ongoing basis.
Still stuck?
The troubleshooting guide covers every artifact we know of, what causes it, and which repairs are worth trying.
Read the troubleshooting guide