The Vocal Remover uses AI audio separation to remove vocals from a song and create vocals plus an
instrumental/karaoke track. You can process audio locally in the browser or use the Free Audio Tools GPU Server
when available, then preview and download the generated WAV files.
Remove vocals and make karaoke tracks in six steps
Load a song, choose local or server processing, start vocal removal, wait for rendering, then preview and
download the vocals or instrumental/karaoke track.
1Upload music file
2Check waveform preview
3Select processing mode
4Remove vocals
5Preview results
6Download WAV or ZIP
Quick start
How to remove vocals from a song
Open the Vocal Remover tool.
Click Upload Audio and select a music file.
Wait for the waveform preview, filename, and duration to appear.
Choose Local processing or Free Audio Tools GPU Server.
Click Remove Vocals.
Wait for the progress indicator to finish.
Preview the generated vocals and instrumental/karaoke track.
Download individual WAV files or the combined ZIP package when available.
Snapshot: upload widget
upload_fileChoose a music file MP3, WAV, OGG, M4A, FLAC, or another browser-supported audio file. Upload Audio
Workspace
Review file metadata and waveform preview
After loading a file, the workspace shows the selected filename, duration, current processing engine, waveform
preview, status text, and progress. Use this area to confirm that the correct source file is loaded before
starting an AI vocal removal job.
Snapshot: vocal remover workspace
Filesong-demo.wav
Duration3:18
EngineLocal processing
Waveform Preview • Ready
Settings
Choose processing mode
The current output mode creates two files: vocals and instrumental/karaoke. Processing Mode controls whether
the AI job runs locally in the browser or on the Free Audio Tools GPU Server.
Local processing depends on the user’s browser, CPU, and memory. GPU Server processing is designed for heavier
workloads but may be temporarily unavailable during cooldown.
Snapshot: vocal removal settings
Processing status
Watch progress while AI vocal removal runs
Vocal removal can take time because the model analyzes the song and reconstructs the vocals and
instrumental/karaoke outputs. The progress area shows a spinner, percentage, and current processing message
while the job is active.
Snapshot: progress indicator
Separating vocals and instrumental... 64%
Keep the tab open until the vocals and karaoke track are ready.
Results
Preview and download generated files
When processing finishes, the result area shows audio players for vocals and instrumental. Use the players to
preview each track, then download individual WAV files. Server jobs may also provide a combined ZIP file.
Choose the mode based on device capability, privacy preference, cooldown availability, and expected processing
time.
Local processingRuns the local MDX model in the browser. It avoids the GPU server cooldown but depends on the user’s CPU, memory, and browser support.
Free Audio Tools GPU ServerUploads the selected file to temporary storage and processes it on the Free Audio Tools GPU Server. This is intended for heavier AI vocal removal jobs.
Cooldown behaviorGPU Server processing is limited between jobs. During cooldown, local processing remains available.
Current output modeThe current public mode generates two tracks: vocals and instrumental/karaoke.
Supported input formats
The tool accepts audio files, but actual decoding depends on what the current browser can read.
MP3Common compressed music format. Browser decoding is generally well supported.
WAVBest option for predictable decoding and uncompressed source quality.
OGGSupported by many modern browsers; support can vary by browser and codec.
M4AUsually supported in Chrome, Edge, and Safari when the codec is browser-decodable.
FLACSupported in many desktop browsers, but compatibility still depends on the browser.
Technical notes
How AI vocal removal works here
Vocal removal does not extract original studio tracks. It estimates likely vocal and instrumental content from
the final mixed song using an AI model trained for source separation.
Client-side previewThe selected file is decoded in the browser to display duration and waveform preview.
Local AI pathLocal processing loads the configured MDX ONNX model and runs inference in the browser when supported.
Server AI pathGPU Server processing sends the audio to the protected server workflow without exposing private API keys to the browser.
Two-track outputThe tool currently returns vocals and instrumental/karaoke WAV files.
Preview linksGenerated files are attached to audio players so users can listen before downloading.
Memory limitsBrowser AI models can fail on low-memory devices or long songs because model inference needs large temporary tensors.
Quality expectations and limitations
AI vocal removal is useful for karaoke practice, remix preparation, cover creation, transcription, and quick
music study. It is not identical to having the original multitrack session. Dense mixes, heavy reverb,
distortion, backing vocals, doubled guitars, cymbals, and instruments sharing the vocal frequency range can
cause bleed or artifacts.
For best results, use the cleanest source file available, avoid low-bitrate files when possible, and process a
full-quality WAV or high-quality MP3 if you have one.
Troubleshooting
The Remove Vocals button is disabledSelect an audio file first. The button enables after the file is loaded and basic audio information is available.
The file does not loadThe browser could not decode that audio format. Try WAV or MP3, or use a current desktop browser.
Local processing is slowAI vocal removal is heavy. Close other tabs, use a shorter file, or choose the GPU Server mode when available.
GPU Server is on cooldownWait for the countdown to finish, or switch Processing Mode to Local processing.
The separated result has artifactsVocal removal is an AI estimate. Dense mixes, reverb, distortion, and overlapping instruments can leave bleed or artifacts.
Audio preview does not playCheck whether the browser blocked autoplay or whether the generated audio URL has expired. Click the play button manually.
No ZIP download appearsThe ZIP button appears only after the server result includes a combined vocals and instrumental package.