Took me longer than I’d like to admit to learn this: the order you clean spoken-word audio in matters as much as the tools you use. Here’s the sequence that finally stopped me from making my own recordings worse.
De-noise before you level. I used to normalize or compress first because the file was too quiet, and every time I did, the HVAC hum and fan noise got louder right along with the voice. Steady background noise is easiest to remove while it’s still low relative to the speech. Once you’ve baked it in with compression, you’re stuck with it.
Then enhance and level. After the noise is gone, do your voice clarity work and hit whatever loudness target your platform expects. This step goes way smoother on a clean file, and you can actually hear what the voice needs.
Be deliberate with filler removal. This is the step people overdo. For a tight solo episode, cutting ums, long silences and mouth sounds tightens the pacing nicely. But on an interview, aggressive removal can gut a guest’s personality. Some people think out loud, and the pauses and stumbles are part of how they talk. I keep filler cuts for moments where they’re genuinely distracting and leave the guest’s natural rhythm alone otherwise.
Compare before and after on every single file. If the cleanup makes the voice sound watery or hollow, back off. A quieter recording with a bit of room tone beats an over-processed one every time.
Finally, squeeze more out of the cleaned file. Transcription is noticeably more accurate on cleaned audio, and from one good transcript you can pull chapters, show notes, summaries and social posts instead of re-listening to the whole episode.
Full disclosure: I built Denoisr, which automates this exact chain (cleanup modes, before/after comparison, transcript and show notes from one upload). But the sequencing logic applies whatever you’re using, Audacity and a DAW included. Curious what order others run their cleanup in.