How This Vocal Remover Actually Works

Updated September 2026

This page walks through the exact math running when you click "Separate vocals" — not a marketing summary, the actual arithmetic on the actual sample data. It's a short explanation because the technique itself is short: this is classic stereo-channel math, not a trained model.

The two lines of math

A stereo audio file is two arrays of numbers — one for the left channel, one for the right — sampled tens of thousands of times per second. For every matching pair of samples L and R, this tool computes:

OutputFormulaWhat it keeps
Instrumentalside = (L − R) / 2Whatever is not identical in both channels — i.e. anything panned off-center
Vocal isolatemid = (L + R) / 2Whatever is identical in both channels — i.e. anything panned dead-center

That's the entire algorithm. There's no frequency analysis, no spectral masking, no trained weights — just subtraction and addition of two number arrays, run once per sample, for every sample in the file.

Why subtraction cancels the vocal

Most commercial mixes place the lead vocal dead-center in the stereo field, meaning the engineer put an identical (or very nearly identical) copy of the vocal signal into both the left and right channels. If a value is genuinely identical in L and R, then L − R = 0 for every sample of that content — it cancels to silence. Anything that's not identical between the channels — a guitar panned left, a synth pad panned right, stereo reverb, doubled harmony vocals spread wide — survives the subtraction, because the two channels disagree on what that content sounds like.

Why addition keeps the vocal

The mirror operation, addition, does the opposite: content that's identical in both channels adds constructively (it gets reinforced, then halved back to the original level by the /2), while content that's genuinely different between channels partially cancels itself out in the sum. So the "vocal" output is really "everything that was mixed the same into both channels" — which is usually dominated by the lead vocal, but not exclusively.

What "center-panned" really means, precisely

This isn't about loudness or how prominent something sounds — it's purely about whether a track's signal is mixed identically into left and right. A kick drum, snare, and bass line are very often panned dead-center too, for the same mixing-convention reasons vocals are (mono compatibility, punch, consistency across playback systems). That's why they show up in the "vocal" output alongside the actual vocal — the math has no way to distinguish "this is a voice" from "this is a kick drum," it only knows "this is centered" versus "this is not."

What actually runs, technically

Your browser's Web Audio API decodes the uploaded file with AudioContext.decodeAudioData — the same built-in decoder your browser already uses to play any mp3 or wav. The two channel arrays are read out with getChannelData(0) and getChannelData(1), the mid/side math above runs in a single loop over plain Float32Arrays, and the two resulting arrays are re-encoded into standard 16-bit PCM WAV files locally using URL.createObjectURL — see the output format page for exactly how that encoding works. None of this touches a network request; there's nothing to upload because there's no server involved in any of it.

Why this is fast and why it has no model to download

Because it's a fixed, tiny amount of arithmetic per sample rather than a neural network's forward pass, a three-minute stereo song (roughly 8 million samples per channel at a typical 44.1kHz) separates in a fraction of a second on ordinary hardware — there's no multi-hundred-megabyte model to download first, and no server-side queue to wait in. The tradeoff for that speed and simplicity is quality: see how this compares to true AI stem separation for the honest version of that tradeoff.

FAQ

Does this analyze frequencies or instruments at all?
No. The math never looks at pitch, timbre, or what instrument is playing — only at whether a sample value matches between the left and right channel. That's both the whole strength (instant, no model, works offline) and the whole limitation (can't tell a vocal from a centered kick drum) of the technique.
Why is the free preview limited to 60 seconds?
That's a product limit on this tool's free tier, unrelated to the math itself — the mid/side calculation runs the same way regardless of length. Full-length processing is a paid tier planned for later; the 60-second preview is free with no limit on how many songs you try.
Is the output really mono, even though it downloads as a stereo WAV?
Yes for each individual output — see the output format page for exactly why it's encoded that way.

→ Try it free