Free · No upload · MP3 & WAV export

Music & song speed changer

Slow down songs to practice without changing the key, speed up lectures, or shift the pitch in semitones without touching the tempo — then download the result.

Drop an audio file here, or browse

MP3, WAV, M4A, OGG… · stays on your device

🔒 Time-stretching and encoding run on your device — your music is never uploaded.

One slider away

🎸

Learn songs faster

Drop a solo to 0.5×–0.75× at full pitch and your ears finally catch every note. The classic practice trick, no app needed.

🎤

Change the key

Backing track too high? Shift it down two semitones and sing comfortably — the tempo doesn't move.

🎧

Lectures & remixes

1.5× a recorded lecture without chipmunk voice, or make nightcore and slowed + pitched edits in two clicks.

Speed, pitch, and separating the two

Playing audio faster by simply reading samples more quickly changes pitch along with speed — the tape-machine effect, where a voice becomes a chipmunk. The two are linked because both are properties of the same waveform in time.

Separating them requires time-stretching, which reconstructs the signal at a new duration while preserving frequency content. The common approaches are phase vocoders, which work in the frequency domain, and overlap-add methods such as WSOLA, which cut the waveform into short segments and repeat or drop them at points chosen to preserve waveform continuity.

Neither is free of artefacts. Phase vocoders can produce a smeared, reverberant quality on transients — drums and plosives suffer most — sometimes described as phasiness. Overlap-add methods can produce a stuttering or warbling texture on sustained tones. Speech generally survives stretching better than music, because it is less harmonically dense.

How far you can push it

Modest changes are essentially transparent. Between about 0.9× and 1.25×, good algorithms produce results most listeners will not identify as processed, which covers most practical use — trimming a podcast, tightening a narration, fitting audio to a video edit.

Beyond about 1.5×, artefacts become noticeable on music and remain acceptable on speech. Speech stays intelligible surprisingly far, and regular listeners at 2× adapt to it, though comprehension of unfamiliar or complex material degrades even where the words are clear. Studies of accelerated speech generally find recall falls before intelligibility does.

Slowing down is harder than speeding up. Stretching to 0.5× requires the algorithm to invent twice as much signal as it was given, and artefacts become pronounced. For transcription work, small decrements around 0.75× are far more useful than dramatic ones.

Deliberate pitch changes

Where you do want pitch to change, the operations are distinct. Pitch shifting without changing duration is the mirror of time-stretching and uses the same algorithms. Changing both together, as a tape machine does, is a resample and is computationally trivial — and it is sometimes exactly the sound wanted, since it is the classic sound of a record played at the wrong speed.

For music, remember that shifting pitch by an arbitrary amount takes it out of standard tuning. Semitone steps preserve musical relationships; a shift of 30 cents leaves everything slightly sharp against anything it is played with. Musical transposition means multiplying frequency by the twelfth root of two per semitone.

Formants are the reason heavily pitch-shifted voices sound artificial. Vocal tract resonances stay roughly fixed as a person changes pitch, but naive shifting moves them with the fundamental, producing the cartoon quality. Formant-preserving algorithms keep them in place and sound markedly more natural on voice.

Speed & pitch FAQ

Why doesn't the voice sound like a chipmunk?

Naively speeding audio up raises the pitch too. This tool uses time-stretching (WSOLA), which repeats or skips tiny grains of audio so tempo and pitch can change independently.

Does extreme stretching sound perfect?

Small changes (0.75×–1.5×, ±4 semitones) sound very clean. Extreme settings introduce mild artifacts on complex music — that's true of any algorithm, including paid apps.

What's the output length?

Duration scales with speed: a 4:00 song at 0.5× becomes 8:00, at 2× it becomes 2:00. Pitch changes don't affect length.

Why does my sped-up audio sound robotic?

Time-stretching artefacts. Phase vocoders smear transients and overlap-add methods can warble on sustained tones. Music suffers more than speech; keeping changes under about 1.5× usually stays clean.

Can I change speed without changing pitch?

Yes — that is time-stretching, which reconstructs the signal at a new duration while preserving frequency content. Simply resampling changes both together, like a tape machine.

How do I slow a song down to learn it?

Load the track and drop the speed to somewhere between 50% and 75%, leaving pitch preservation on so the key does not fall with the tempo — otherwise you learn the part in the wrong key. Work at a speed where you can play it cleanly, then move up in small steps. If you want to practise against a different key, use the semitone control instead, which shifts pitch and leaves the tempo alone.