How VocaSplitter works
The whole process is one upload and one choice, but a few decisions along the way make a large difference to how usable the result is. This page walks through each step and explains what is happening underneath.
In short: upload the best quality version of the song you have, choose the fewest stems that answer your question, preview on headphones, and download during the session — files are deleted afterwards.
What separation actually does
It helps to know what is and is not possible before you start. When a song was mixed down, every instrument was summed into two channels, and the information about which sound came from where was permanently lost. Separation does not undo that — nothing can.
What the model does is estimate, for every tiny slice of time and frequency in the recording, how much of that energy belongs to each instrument, and then build new audio files from those estimates. That single idea explains almost every quirk you will encounter: why sparse recordings separate beautifully, why a cymbal sometimes appears in the vocal stem, and why a very loud master leaves a faint ghost of the singer behind.
If you want the full picture, including how the models are trained and why the artifacts sound the way they do, read how AI stem separation actually works.
Step by step
- 1
Pick the best version of the song you have
Source quality decides more than any setting. A CD rip, a purchased WAV or a FLAC gives the model full bandwidth to work with. A 128 kbps MP3 has already had most content above roughly 16 kHz stripped by the encoder, and the separation will faithfully reproduce a song with no top end. If all you have is a low-bitrate file, the tool will still work — just expect a duller result and more vocal residue.
- 2
Upload the file
Drag it onto the upload area on the home page, or click to browse. MP3, WAV, M4A, AAC, OGG, FLAC and AIFF are all accepted. Nothing else is required — no account, no email. The browser reads the file locally first to draw the waveform, then uploads it for processing.
- 3
Check the waveform and duration
The preview screen shows the waveform and the length of what you uploaded, which is a quick way to catch the wrong file or a truncated download before spending processing time on it. If the waveform looks like a solid block with no dynamics, the track is heavily limited and will be harder to separate cleanly.
- 4
Choose two, three or five stems
Two stems splits into vocals and instrumental and is the right choice for karaoke tracks and acapellas. Three adds a separate drum track, which suits beat transcription and DJ edits. Five gives vocals, drums, bass, piano and other, and is for when you need the individual parts. Choose the fewest that answer your question — each extra class is another decision the model can get wrong.
- 5
Start the split and wait
Processing runs on our server, not in your browser, so you are not limited by your device. A three-minute song usually takes one to three minutes depending on queue load and mode. Keep the tab open — the page polls for completion and shows your results when the job finishes.
- 6
Preview each stem before you commit
Every stem has its own play control. Listen to the ones that matter for your project, and specifically check the sparsest verse and the loudest chorus — those are where problems show up first. Use headphones; laptop speakers will not reproduce the range where separation artifacts live.
- 7
Download what you need
Download stems individually, or take the whole set as a single ZIP. Do this during the session: files are deleted after processing and the links are short-lived, so a bookmark will not work tomorrow. Save the original file alongside the stems if the project matters — you will want it as a reference.
Advertisement
Getting better results
A handful of habits account for most of the difference between a usable stem and a disappointing one.
- Use the least compressed master you can find. A mix crushed to a very high perceived loudness leaves the model less contrast to work with. Where a pre-master or a CD rip is available, it will separate better than the streaming version.
- Do not re-separate a stem. Feeding a stem back through the tool treats its artifacts as signal and compounds them. Always go back to the original mix and change the mode.
- Judge on headphones, in the sparse sections. The first verse of a song is where residue is most audible, because there is least else going on to mask it.
- Keep the original file with your stems. It is the reference you will want constantly — for checking whether an artifact came from separation or was on the record all along.
- Accept when a song will not cooperate. Some mixes simply do not give up a clean instrumental. Choosing a different track is a better use of an afternoon than fighting a hopeless one.
What to do next
Make a karaoke track
Levels, ghost removal and the mistakes that ruin a backing track on a real PA.
Import stems into a DAW
Sample rates, alignment, tempo mapping and gain staging without losing sync.
Fix an artifact
What each problem is, and which repairs help rather than hurt.
Browse the FAQ
Formats, limits, privacy and rights, answered briefly.
Ready to try it
Upload a track and see how it separates. It costs nothing and takes a couple of minutes.
Open the separator