Free · Automatic · Nothing uploaded

Remove silence from audio

Automatically detect and cut silent sections from podcasts, voice memos and recordings. Tune the threshold, preview every cut on the waveform, and export MP3 or WAV — all on your device.

🤫

Drop an audio file here or click to browse

MP3 · WAV · M4A · OGG — decoded on your device

🔒 Detection and cutting run entirely in your browser — your recording never leaves your device.

How to remove silence from a recording

🎙️

1. Load your audio

Drop in an MP3, WAV, M4A or any audio your browser can play. It decodes locally.

🎚️

2. Tune the detection

Set how quiet counts as silence (threshold) and how long a pause must be before it is cut.

✂️

3. Preview & export

Red regions show what will be removed. Preview the result, then export MP3 or WAV.

Long pauses make podcasts and lecture recordings feel slow — trimming them can shorten a recording by 20–40% without losing a word. The remover keeps a little padding around speech so cuts never sound abrupt, and joins every cut with a micro crossfade to avoid clicks. For manual precision cuts use the Audio Cutter, or level the volume afterwards with the Audio Normalizer.

What counts as silence

Silence in a recording is almost never digital zero. A room has an ambient noise floor — ventilation, traffic, electrical hum, the microphone's own self-noise — so detection works on a threshold rather than on absolute silence. Anything below the threshold for longer than a minimum duration is treated as a gap.

Setting the threshold is the whole problem. Too high and quiet speech is cut, clipping the ends of sentences and removing soft consonants. Too low and nothing is detected, because the noise floor never falls beneath it. A threshold around 6 to 10 dB above the measured noise floor is a reasonable starting point, and the noise floor should be measured from a passage with no speech.

The minimum duration matters as much. Natural speech contains short pauses between words and stop consonants that are genuinely silent for tens of milliseconds — a minimum gap length under about 200 milliseconds will start cutting inside words, producing speech that sounds clipped and breathless.

Removing pauses without destroying rhythm

Removing every pause produces speech that is technically continuous and exhausting to listen to. Pauses carry meaning: they mark clause boundaries, signal emphasis, and give a listener time to process. Speech with all pauses removed feels relentless, and comprehension falls even though nothing was lost.

The more useful operation is shortening rather than eliminating — reducing a three-second pause to half a second, rather than to nothing. Most editing tools that do this well work on a target pause length rather than on removal, and the result sounds tightened rather than compressed.

Breaths are a separate decision. Removing them entirely is a recognisable style and sounds slightly unnatural; reducing their level rather than cutting them is generally better, since a breath before a sentence is part of how speech is understood. Cutting a breath also removes the room tone under it, which can leave an audible hole.

Avoiding audible edits

Every cut is a potential click, for the same reason as any other edit: an instantaneous jump in sample value. Short fades at each boundary — a handful of milliseconds — remove this and are imperceptible.

The subtler artefact is the loss of room tone. Cutting silence out of a recording removes the ambient bed along with it, so the background noise stops abruptly at each edit and resumes when speech returns. In a noticeably noisy recording this pumping is more distracting than the pauses were.

Two approaches address it: attenuate the silent passages rather than deleting them, keeping duration and reducing level, or delete them and lay a continuous bed of room tone underneath the whole programme. Recording thirty seconds of room tone at every session — standard practice in audio production — is what makes the second option available. If levels also vary across the recording, handle that with the normaliser before cutting rather than after.

Silence remover FAQ

How does silence detection work?

The tool measures loudness in short windows across the whole file. Stretches quieter than your threshold that last longer than the minimum duration are marked as silence and removed, with padding kept around speech so words are never clipped.

What threshold should I use?

-40 dB works for most voice recordings. If too much is being cut, lower it towards -50 dB; if breaths and room noise survive, raise it towards -30 dB. The waveform shows exactly what will be removed before you commit.

Will the cuts be audible?

No — every join uses a short crossfade, and the padding setting keeps a natural breath of air around each phrase.

Is my audio uploaded?

No. Decoding, detection and export all run inside your browser. Nothing is sent to any server.

Why is the tool cutting the ends of my words?

The threshold is too high or the minimum gap too short. Soft consonants and trailing syllables fall below aggressive thresholds — set the threshold around 6 to 10 dB above the measured noise floor and keep the minimum gap above 200 milliseconds.

Why does the background noise pulse after editing?

Cutting silence removes the room tone with it, so ambience stops and starts at each edit. Attenuate the pauses instead of deleting them, or lay continuous room tone under the whole programme.