Free Audio Tools

Tool guide

Vocal Remover documentation

The Vocal Remover uses AI audio separation to remove vocals from a song and create vocals plus an instrumental/karaoke track. You can process audio locally in the browser or use the Free Audio Tools GPU Server when available, then preview and download the generated WAV files.

Workflow overview

Remove vocals and make karaoke tracks in six steps

Load a song, choose local or server processing, start vocal removal, wait for rendering, then preview and download the vocals or instrumental/karaoke track.

1Upload music file
2Check waveform preview
3Select processing mode
4Remove vocals
5Preview results
6Download WAV or ZIP

Quick start

How to remove vocals from a song

  1. Open the Vocal Remover tool.
  2. Click Upload Audio and select a music file.
  3. Wait for the waveform preview, filename, and duration to appear.
  4. Choose Local processing or Free Audio Tools GPU Server.
  5. Click Remove Vocals.
  6. Wait for the progress indicator to finish.
  7. Preview the generated vocals and instrumental/karaoke track.
  8. Download individual WAV files or the combined ZIP package when available.
Snapshot: upload widget
Choose a music file MP3, WAV, OGG, M4A, FLAC, or another browser-supported audio file. Upload Audio

Workspace

Review file metadata and waveform preview

After loading a file, the workspace shows the selected filename, duration, current processing engine, waveform preview, status text, and progress. Use this area to confirm that the correct source file is loaded before starting an AI vocal removal job.

Snapshot: vocal remover workspace
File song-demo.wav
Duration 3:18
Engine Local processing

Waveform Preview • Ready

Settings

Choose processing mode

The current output mode creates two files: vocals and instrumental/karaoke. Processing Mode controls whether the AI job runs locally in the browser or on the Free Audio Tools GPU Server.

Local processing depends on the user’s browser, CPU, and memory. GPU Server processing is designed for heavier workloads but may be temporarily unavailable during cooldown.
Snapshot: vocal removal settings

Processing status

Watch progress while AI vocal removal runs

Vocal removal can take time because the model analyzes the song and reconstructs the vocals and instrumental/karaoke outputs. The progress area shows a spinner, percentage, and current processing message while the job is active.

Snapshot: progress indicator
Separating vocals and instrumental... 64%

Keep the tab open until the vocals and karaoke track are ready.

Results

Preview and download generated files

When processing finishes, the result area shows audio players for vocals and instrumental. Use the players to preview each track, then download individual WAV files. Server jobs may also provide a combined ZIP file.

Snapshot: vocal remover results

Vocals

Download WAV

Instrumental / Karaoke

Download WAV

Vocals + Instrumental ZIP

Download vocals and instrumental together.

Download ZIP

Processing modes

Choose the mode based on device capability, privacy preference, cooldown availability, and expected processing time.

Local processing Runs the local MDX model in the browser. It avoids the GPU server cooldown but depends on the user’s CPU, memory, and browser support.
Free Audio Tools GPU Server Uploads the selected file to temporary storage and processes it on the Free Audio Tools GPU Server. This is intended for heavier AI vocal removal jobs.
Cooldown behavior GPU Server processing is limited between jobs. During cooldown, local processing remains available.
Current output mode The current public mode generates two tracks: vocals and instrumental/karaoke.

Supported input formats

The tool accepts audio files, but actual decoding depends on what the current browser can read.

MP3 Common compressed music format. Browser decoding is generally well supported.
WAV Best option for predictable decoding and uncompressed source quality.
OGG Supported by many modern browsers; support can vary by browser and codec.
M4A Usually supported in Chrome, Edge, and Safari when the codec is browser-decodable.
FLAC Supported in many desktop browsers, but compatibility still depends on the browser.

Technical notes

How AI vocal removal works here

Vocal removal does not extract original studio tracks. It estimates likely vocal and instrumental content from the final mixed song using an AI model trained for source separation.

Client-side preview The selected file is decoded in the browser to display duration and waveform preview.
Local AI path Local processing loads the configured MDX ONNX model and runs inference in the browser when supported.
Server AI path GPU Server processing sends the audio to the protected server workflow without exposing private API keys to the browser.
Two-track output The tool currently returns vocals and instrumental/karaoke WAV files.
Preview links Generated files are attached to audio players so users can listen before downloading.
Memory limits Browser AI models can fail on low-memory devices or long songs because model inference needs large temporary tensors.

Quality expectations and limitations

AI vocal removal is useful for karaoke practice, remix preparation, cover creation, transcription, and quick music study. It is not identical to having the original multitrack session. Dense mixes, heavy reverb, distortion, backing vocals, doubled guitars, cymbals, and instruments sharing the vocal frequency range can cause bleed or artifacts.

For best results, use the cleanest source file available, avoid low-bitrate files when possible, and process a full-quality WAV or high-quality MP3 if you have one.

Troubleshooting

The Remove Vocals button is disabled Select an audio file first. The button enables after the file is loaded and basic audio information is available.
The file does not load The browser could not decode that audio format. Try WAV or MP3, or use a current desktop browser.
Local processing is slow AI vocal removal is heavy. Close other tabs, use a shorter file, or choose the GPU Server mode when available.
GPU Server is on cooldown Wait for the countdown to finish, or switch Processing Mode to Local processing.
The separated result has artifacts Vocal removal is an AI estimate. Dense mixes, reverb, distortion, and overlapping instruments can leave bleed or artifacts.
Audio preview does not play Check whether the browser blocked autoplay or whether the generated audio URL has expired. Click the play button manually.
No ZIP download appears The ZIP button appears only after the server result includes a combined vocals and instrumental package.