Formats must match before joining
Joining audio requires that the files share sample rate, bit depth and channel count. Concatenating a 44.1 kHz file with a 48 kHz one without resampling produces the second at the wrong speed and pitch — about nine percent fast — which is unmistakable and a common cause of a merge sounding wrong.
Channel count must match too. Appending a mono file to a stereo one produces either an error or a section playing from one side. Converting mono to stereo by duplicating the channel is straightforward; converting stereo to mono means summing, which can cause phase cancellation where the two channels contain out-of-phase content.
This tool normalises the parameters before joining, which is why an output may differ in format from one of its inputs. If a merged file sounds slowed, sped up or one-sided, mismatched source parameters are the first thing to check.
Level and tone consistency across sources
The most common flaw in a merged file is not a bad join but inconsistent loudness between segments. Recordings from different sessions, devices or speakers differ in level, and a listener adjusting volume for one section finds the next too loud or too quiet.
Match loudness before joining, using a loudness measurement rather than peak levels — two segments peaking identically can differ substantially in perceived loudness. The normaliser handles this, and doing it before the merge is far easier than fixing it afterwards.
Background noise is the harder mismatch. A cut between a segment with a quiet room tone and one with audible hiss draws attention to itself at every transition. A short crossfade smooths the change, and adding a low level of consistent room tone under the whole programme is the standard trick for disguising it entirely.
Joins, gaps and gapless playback
A direct butt join between two files usually needs a short crossfade to avoid a click, for the same reason cuts do — an instantaneous jump in sample value produces a broadband transient. A few milliseconds suffices where the material is continuous.
Deliberate gaps are a different decision. Concatenating spoken segments with no pause sounds rushed and unnatural; a pause of 300 to 700 milliseconds between topics reads as a natural breath, and around a second signals a section change. Silence in a recording is rarely digital silence, so inserting true zero samples between segments recorded in a room is audible as a hole — better to insert a segment of the room's own quiet tone.
For music, gapless playback matters. Albums mixed to run continuously require that no encoder padding is inserted between tracks, and lossy formats add exactly that unless the encoder writes gapless metadata. This is why a continuously mixed album can develop small gaps after conversion, and why a single merged file is often the safer delivery for a continuous mix.