What counts as silence
Silence in a recording is almost never digital zero. A room has an ambient noise floor — ventilation, traffic, electrical hum, the microphone's own self-noise — so detection works on a threshold rather than on absolute silence. Anything below the threshold for longer than a minimum duration is treated as a gap.
Setting the threshold is the whole problem. Too high and quiet speech is cut, clipping the ends of sentences and removing soft consonants. Too low and nothing is detected, because the noise floor never falls beneath it. A threshold around 6 to 10 dB above the measured noise floor is a reasonable starting point, and the noise floor should be measured from a passage with no speech.
The minimum duration matters as much. Natural speech contains short pauses between words and stop consonants that are genuinely silent for tens of milliseconds — a minimum gap length under about 200 milliseconds will start cutting inside words, producing speech that sounds clipped and breathless.
Removing pauses without destroying rhythm
Removing every pause produces speech that is technically continuous and exhausting to listen to. Pauses carry meaning: they mark clause boundaries, signal emphasis, and give a listener time to process. Speech with all pauses removed feels relentless, and comprehension falls even though nothing was lost.
The more useful operation is shortening rather than eliminating — reducing a three-second pause to half a second, rather than to nothing. Most editing tools that do this well work on a target pause length rather than on removal, and the result sounds tightened rather than compressed.
Breaths are a separate decision. Removing them entirely is a recognisable style and sounds slightly unnatural; reducing their level rather than cutting them is generally better, since a breath before a sentence is part of how speech is understood. Cutting a breath also removes the room tone under it, which can leave an audible hole.
Avoiding audible edits
Every cut is a potential click, for the same reason as any other edit: an instantaneous jump in sample value. Short fades at each boundary — a handful of milliseconds — remove this and are imperceptible.
The subtler artefact is the loss of room tone. Cutting silence out of a recording removes the ambient bed along with it, so the background noise stops abruptly at each edit and resumes when speech returns. In a noticeably noisy recording this pumping is more distracting than the pauses were.
Two approaches address it: attenuate the silent passages rather than deleting them, keeping duration and reducing level, or delete them and lay a continuous bed of room tone underneath the whole programme. Recording thirty seconds of room tone at every session — standard practice in audio production — is what makes the second option available. If levels also vary across the recording, handle that with the normaliser before cutting rather than after.