Home / Vocal Remover
Vocal Remover
Remove vocals from any song to get karaoke instrumentals or isolated vocals. Powered by Demucs (Meta AI).
Upload a Song
Upload the audio track you want to separate into vocals and instrumentals.
Drag & drop your file here, or browse
Supports MP3, WAV, FLAC, OGG, M4A. Max 50MB. Use highest quality source for best results.file.mp3
0 MB
Separating vocals from instrumentals... This typically takes 30-90 seconds.
The song with vocals removed -- perfect for karaoke or remixing.
0:00
0:00
The isolated vocal track -- useful for acapella, sampling, or lyric transcription.
0:00
0:00
Credit Cost
5 credits (flat rate per track)
Fixed cost regardless of audio duration.
What You Get
- Instrumental track -- the song without vocals
- Vocal track -- isolated vocals only
Tips for Best Results
- Use the highest quality source file (WAV/FLAC preferred)
- Avoid heavily compressed or low-bitrate MP3 files
- Studio recordings produce cleaner separation
Best For
Karaoke creation, remix preparation, acapella extraction, vocal sampling, and lyric transcription.
Frequently Asked Questions
AI vocal removal uses deep learning models (Demucs by Meta AI) to analyze the frequency patterns and spatial positioning of audio elements. It identifies and separates the vocal track from the instrumental track, producing clean karaoke-style output.
Demucs v4 produces state-of-the-art vocal separation quality. The instrumental output is typically clean enough for karaoke or remixing. Some bleed-through may occur on heavily processed or compressed tracks.
Yes! You get both the isolated vocals and the instrumental track. This is useful for creating acapella versions, sampling vocals for remixes, or transcribing lyrics from complex mixes.
We support MP3, WAV, FLAC, OGG, M4A, and WEBM files up to 50MB. For best results, use the highest quality source file available (WAV or FLAC preferred over compressed MP3).
Vocal removal uses the stem splitting pipeline and costs 5 credits flat per file processed. You receive both the isolated vocals and the instrumental track. Free users can try the tool with starter credits after signing up for an account.
Processing typically takes 15-60 seconds depending on the file length and quality. A standard 3-4 minute MP3 song usually completes in about 20 seconds. Longer tracks or lossless formats may take slightly more time.
Yes. The instrumental output is specifically designed for karaoke use. The vocal track is removed cleanly enough for live karaoke performance. For best results, start with a high-quality source file. Some vocal artifacts may remain on heavily produced tracks.
Separation quality depends on the source material. Songs with clear vocal/instrumental separation tend to produce the best results. Heavily processed vocals, vocal harmonies layered with instruments, or low-quality source files may result in some bleed-through. Try using a higher quality source file for better results.
Yes, but results vary. Live recordings with close-miked vocals work well. Ambient live recordings where the vocals blend with the room sound and audience noise will have lower separation quality. Consider running the Audio Enhancer first to clean up the recording.
Yes. Upload your audio file to the vocal removal API endpoint with your API key. The API returns two files: the isolated vocal track and the instrumental track. This enables batch processing for music production, karaoke services, and content creation workflows.
Yes. The isolated vocal track can be imported into any DAW (Ableton, FL Studio, Logic Pro, etc.) for remixing, sampling, or acapella arrangements. The extracted vocals maintain the original timing and pitch, making them easy to align with new instrumentals.
Demucs is a state-of-the-art music source separation model developed by Meta AI Research. It uses a hybrid transformer architecture trained on thousands of songs to achieve best-in-class separation quality. WhatIsTTS routes vocal removal requests to Demucs via our provider gateway for consistent, high-quality results.
Need a specific TTS workflow?
Compare providers, test voices, then run it through one brokered API.