Updated September 2026
This page walks through the exact math running when you click "Separate vocals" — not a marketing summary, the actual arithmetic on the actual sample data. It's a short explanation because the technique itself is short: this is classic stereo-channel math, not a trained model.
A stereo audio file is two arrays of numbers — one for the left channel, one for the right — sampled tens of thousands of times per second. For every matching pair of samples L and R, this tool computes:
| Output | Formula | What it keeps |
|---|---|---|
| Instrumental | side = (L − R) / 2 | Whatever is not identical in both channels — i.e. anything panned off-center |
| Vocal isolate | mid = (L + R) / 2 | Whatever is identical in both channels — i.e. anything panned dead-center |
That's the entire algorithm. There's no frequency analysis, no spectral masking, no trained weights — just subtraction and addition of two number arrays, run once per sample, for every sample in the file.
Most commercial mixes place the lead vocal dead-center in the stereo field, meaning the engineer put an identical (or very nearly identical) copy of the vocal signal into both the left and right channels. If a value is genuinely identical in L and R, then L − R = 0 for every sample of that content — it cancels to silence. Anything that's not identical between the channels — a guitar panned left, a synth pad panned right, stereo reverb, doubled harmony vocals spread wide — survives the subtraction, because the two channels disagree on what that content sounds like.
The mirror operation, addition, does the opposite: content that's identical in both channels adds constructively (it gets reinforced, then halved back to the original level by the /2), while content that's genuinely different between channels partially cancels itself out in the sum. So the "vocal" output is really "everything that was mixed the same into both channels" — which is usually dominated by the lead vocal, but not exclusively.
This isn't about loudness or how prominent something sounds — it's purely about whether a track's signal is mixed identically into left and right. A kick drum, snare, and bass line are very often panned dead-center too, for the same mixing-convention reasons vocals are (mono compatibility, punch, consistency across playback systems). That's why they show up in the "vocal" output alongside the actual vocal — the math has no way to distinguish "this is a voice" from "this is a kick drum," it only knows "this is centered" versus "this is not."
Your browser's Web Audio API decodes the uploaded file with AudioContext.decodeAudioData — the same built-in decoder your browser already uses to play any mp3 or wav. The two channel arrays are read out with getChannelData(0) and getChannelData(1), the mid/side math above runs in a single loop over plain Float32Arrays, and the two resulting arrays are re-encoded into standard 16-bit PCM WAV files locally using URL.createObjectURL — see the output format page for exactly how that encoding works. None of this touches a network request; there's nothing to upload because there's no server involved in any of it.
Because it's a fixed, tiny amount of arithmetic per sample rather than a neural network's forward pass, a three-minute stereo song (roughly 8 million samples per channel at a typical 44.1kHz) separates in a fraction of a second on ordinary hardware — there's no multi-hundred-megabyte model to download first, and no server-side queue to wait in. The tradeoff for that speed and simplicity is quality: see how this compares to true AI stem separation for the honest version of that tradeoff.