- Home
- Claude video editing
- Remove filler words
On this page
What you need
- Claude Code in a terminal, pointed at a folder with your footage.
- ffmpeg and ffprobe on your PATH.
- A local Whisper model for word-level transcripts.
- Python with numpy, soundfile and OpenCV for the checks.
Why normal transcripts miss the ums
Whisper and most speech-to-text models are trained to write clean text, so they drop "um", "uh" and stutters. Prime the model with a prompt full of fillers and it writes them down, with timestamps.
That is the step every "remove filler words" button hides. Descript, Captions and the rest run their own detector. In Claude Code you run the detector yourself, which means you can see it, tune it and check it.
import whisper, json
model = whisper.load_model("turbo")
r = model.transcribe("work/episode_16k.wav", word_timestamps=True,
condition_on_previous_text=False,
initial_prompt="Umm, so, uh, I was like, uhh, you know, hmm... Um, yeah. Uh, we, um, like, uhm, basically, uh, okay.")
json.dump(r, open("work/words.json", "w"))The tightening pass, step by step
Transcribe primed, mark every filler, stutter and pause over 0.4 s, cut in energy valleys, fade each join, render, then transcribe the output to prove it is clean.
Extract the audio
A 16 kHz mono wav is all Whisper needs. Your camera file stays read-only.
Transcribe with the filler prime
Word timestamps on, previous-text conditioning off, and an initial prompt written in the voice of someone who says "um" a lot.
Mark what goes
Filler sounds (um, uh, er), filler words that carry no meaning (you know, I mean, basically), stutters ("the the"), false starts, restatements, and pauses over 0.4 s. Keep "like" when it means something: "tools like ffmpeg" stays.
Cut in the valleys
Only cut where the audio drops at least 5 dB below the speech around it. Transcript word ends run 100 to 300 ms early, so a cut placed on the timestamp clips the last consonant.
Render with fades
10 to 20 ms fades on every audio edit. Ranges that touch in the source get merged, so the same sound never plays twice.
Prove it
Transcribe the finished file with the same prime and count the fillers left. Listen to every join. If an edit cannot be proven clean, leave the filler in.
ffmpeg -i episode.mp4 -vn -ac 1 -ar 16000 work/episode_16k.wav[
{ "start": 3091.42, "end": 3094.87, "why": "kept" },
{ "start": 3096.12, "end": 3099.40, "why": "kept, 'like' and a restatement removed" },
{ "start": 3099.62, "end": 3104.05, "why": "kept, 0.9 s pause trimmed to 0.2 s" }
]ffmpeg -i episode.mp4 -filter_complex "\
[0:v]trim=3091.42:3094.87,setpts=PTS-STARTPTS[v0];\
[0:a]atrim=3091.42:3094.87,asetpts=PTS-STARTPTS,afade=t=out:st=3.435:d=0.015[a0];\
[0:v]trim=3096.12:3099.40,setpts=PTS-STARTPTS[v1];\
[0:a]atrim=3096.12:3099.40,asetpts=PTS-STARTPTS,afade=t=in:d=0.015[a1];\
[v0][a0][v1][a1]concat=n=2:v=1:a=1[v][a]" -map "[v]" -map "[a]" out/clip.mp4What to check before you post
Three things break a filler edit: a clipped word, a doubled word at a join, and a filler the transcript invented. Each has a check.
- Clipped words. Check the word before every cut is complete to its last letter ("customers", not "customer"). Extend the edge to keep the whole final sound, but stop before the next word starts, or you get a stray hiss.
- Doubled audio. Correlate the half second after each join with the half second before it. A match means the same source played twice ("ad spend spend").
- Invented words. Transcription sometimes hears "So..." or "Because..." at a cut that is not in the audio. Confirm a suspected leak a second way before cutting deeper into a real word.
- Hidden fillers. A flat, steady voiced stretch of 0.2 to 0.5 s before a pause is an "um" even when no transcript wrote it down.
words = [w for s in r["segments"] for w in s["words"]]
left = [(w["word"], round(w["start"], 1)) for w in words
if w["word"].strip(" .,!?").lower() in ("um", "uh", "umm", "uhh", "er")]
print(len(left), "fillers left", left[:6])Every Claude video editing guide
Frequently asked questions
How do I remove filler words from a video automatically?
Transcribe it with word timestamps and a prompt full of "um, uh, like" so the fillers are written down, cut each filler in the quiet gap between words, fade every join by 10 to 20 ms, and re-transcribe the result to confirm none are left. Claude Code runs all of it with ffmpeg and Whisper.
Can Claude remove ums from audio?
Yes. Claude Code reads a primed Whisper transcript of the audio, writes a keep list of every range between the fillers, and renders it with ffmpeg. It works the same on a wav, an mp3 or the audio of a video.
Why does my transcript not show um and uh?
Speech-to-text models are trained to produce clean text, so they leave disfluencies out. Passing an initial prompt that contains fillers makes Whisper transcribe them.
Will cutting fillers make the video look jumpy?
On one fixed camera, yes, unless you hide the jump. Use a camera change, a short punch-in, or a 5-frame dissolve on horizontal edits. Vertical shorts tolerate harder jump cuts.
Want finished ads, not editing sessions?
Director by Sprites is a full creative team for your startup. Every Monday you get finished, on-brand video ads, reviewed by a human. Strategy, hooks, scripts, shooting and editing, end to end. You never write a prompt. You pick the winners. Sprites ships the campaigns. In private beta.