Vocal Remover: Karaoke & Acapella Maker Online
Split a song into vocals and instrumental with an on-device separation model. Karaoke and acapella tracks from one pass.
Under the hood this is Kim_Vocal_2, a community MDX-Net model distributed through the Ultimate Vocal Remover project and credited to its trainer KimberleyJSN. It reads the mix as a 7680-point spectrogram in overlapping 256-frame windows and predicts the vocal spectrum directly; subtracting that from the mix yields the instrumental. That subtraction is why the two stems always sum back to the original. Nothing is invented, only attributed. Expect clean karaoke tracks from typical produced music, with the usual caveats of source separation: heavy reverb tails and crowd vocals blur the boundary, and the model works on the first six minutes of very long files. Both stems export as WAV; pick which one the download should be with the Output stem setting.
How to Use It
- Upload your audio file and wait for the local preview to load.
- Adjust the settings: Pick which stem the download should be, instrumental for karaoke or vocal for an acapella. And start the separation; a full song takes minutes on-device.
- Download the result: Download the processed file directly to your device.
Controls and Options
Instrumental is the default and covers karaoke, remix beds and practice tracks. Switch to Vocal for acapellas, samples and vocal studies. The two stems always sum back to the original mix, so processing twice gives you a complete pair.
Formats and Output
Reads common audio formats your browser can decode, including MP3, WAV, M4A, OGG, and FLAC, and exports the chosen stem as a lossless 16-bit WAV.
Frequently Asked Questions
How good is the separation?
Kim_Vocal_2 is one of the most widely used community vocal models, and the shipped engine is parity-locked to its reference implementation. Produced studio music separates cleanly; live recordings, heavy reverb and background choirs are harder for any separation model.
Why does it take minutes?
The model analyzes the full spectrum of the song in overlapping windows on your device. A several-minute song means hundreds of model passes; nothing is being uploaded or queued, the time is the computation itself.
How do I get both stems?
Run the separation with Instrumental selected, download, switch the Output stem setting to Vocal and process again. The two files sum back to the original mix.
What input works best?
A produced stereo mix at CD quality or better. Very low-bitrate sources carry compression artifacts into both stems, and mono recordings give the model less spatial information to work with.
Is there a length limit?
The first six minutes of a file are separated in one pass; anything beyond that is left out and the result text says so. For an album-length file, cut the song out first with Audio Cutter.
Is my audio uploaded anywhere?
Supported media processing runs on your device and file contents are not uploaded for processing. Site analytics and the assistant are separate.