Workflow · Updated October 2, 2026

Remove filler wordsfrom video with Claude

Claude Code removes ums, uhs, stutters and dead air by reading a word-level transcript primed to keep fillers, then cutting each one in the quiet gap between words with ffmpeg. A plain transcript hides most fillers, so the prime is the whole trick. Every cut is checked by transcribing the result again.

Edit the transcript, Claude cuts the timeline

theyruntheiruseracquisition,like,paidadsandthatsideofthebusiness,0.9 salmostlikeahedgefund

Keptthey run their user acquisition almost like a hedge fund

  • filler from a primed transcript
  • restatement, cut in the energy valley
  • pause over 0.4 s, trimmed to 0.2 s

What you need

  • Claude Code in a terminal, pointed at a folder with your footage.
  • ffmpeg and ffprobe on your PATH.
  • A local Whisper model for word-level transcripts.
  • Python with numpy, soundfile and OpenCV for the checks.

Why normal transcripts miss the ums

Whisper and most speech-to-text models are trained to write clean text, so they drop "um", "uh" and stutters. Prime the model with a prompt full of fillers and it writes them down, with timestamps.

That is the step every "remove filler words" button hides. Descript, Captions and the rest run their own detector. In Claude Code you run the detector yourself, which means you can see it, tune it and check it.

Word-level transcript, primed so the fillers survive.
import whisper, json
model = whisper.load_model("turbo")
r = model.transcribe("work/episode_16k.wav", word_timestamps=True,
                     condition_on_previous_text=False,
                     initial_prompt="Umm, so, uh, I was like, uhh, you know, hmm... Um, yeah. Uh, we, um, like, uhm, basically, uh, okay.")
json.dump(r, open("work/words.json", "w"))

The tightening pass, step by step

Transcribe primed, mark every filler, stutter and pause over 0.4 s, cut in energy valleys, fade each join, render, then transcribe the output to prove it is clean.

  1. Extract the audio

    A 16 kHz mono wav is all Whisper needs. Your camera file stays read-only.

  2. Transcribe with the filler prime

    Word timestamps on, previous-text conditioning off, and an initial prompt written in the voice of someone who says "um" a lot.

  3. Mark what goes

    Filler sounds (um, uh, er), filler words that carry no meaning (you know, I mean, basically), stutters ("the the"), false starts, restatements, and pauses over 0.4 s. Keep "like" when it means something: "tools like ffmpeg" stays.

  4. Cut in the valleys

    Only cut where the audio drops at least 5 dB below the speech around it. Transcript word ends run 100 to 300 ms early, so a cut placed on the timestamp clips the last consonant.

  5. Render with fades

    10 to 20 ms fades on every audio edit. Ranges that touch in the source get merged, so the same sound never plays twice.

  6. Prove it

    Transcribe the finished file with the same prime and count the fillers left. Listen to every join. If an edit cannot be proven clean, leave the filler in.

Pull a 16 kHz mono track for transcription. The original file is never touched.
ffmpeg -i episode.mp4 -vn -ac 1 -ar 16000 work/episode_16k.wav
The edit list Claude writes: source ranges to keep, in seconds. Nothing else gets cut.
[
  { "start": 3091.42, "end": 3094.87, "why": "kept" },
  { "start": 3096.12, "end": 3099.40, "why": "kept, 'like' and a restatement removed" },
  { "start": 3099.62, "end": 3104.05, "why": "kept, 0.9 s pause trimmed to 0.2 s" }
]
Render two kept ranges with a 15 ms audio fade at the join. Claude generates this for every range in the list.
ffmpeg -i episode.mp4 -filter_complex "\
[0:v]trim=3091.42:3094.87,setpts=PTS-STARTPTS[v0];\
[0:a]atrim=3091.42:3094.87,asetpts=PTS-STARTPTS,afade=t=out:st=3.435:d=0.015[a0];\
[0:v]trim=3096.12:3099.40,setpts=PTS-STARTPTS[v1];\
[0:a]atrim=3096.12:3099.40,asetpts=PTS-STARTPTS,afade=t=in:d=0.015[a1];\
[v0][a0][v1][a1]concat=n=2:v=1:a=1[v][a]" -map "[v]" -map "[a]" out/clip.mp4

What to check before you post

Three things break a filler edit: a clipped word, a doubled word at a join, and a filler the transcript invented. Each has a check.

  • Clipped words. Check the word before every cut is complete to its last letter ("customers", not "customer"). Extend the edge to keep the whole final sound, but stop before the next word starts, or you get a stray hiss.
  • Doubled audio. Correlate the half second after each join with the half second before it. A match means the same source played twice ("ad spend spend").
  • Invented words. Transcription sometimes hears "So..." or "Because..." at a cut that is not in the audio. Confirm a suspected leak a second way before cutting deeper into a real word.
  • Hidden fillers. A flat, steady voiced stretch of 0.2 to 0.5 s before a pause is an "um" even when no transcript wrote it down.
The last check on a finished file: transcribe it primed and count what is left.
words = [w for s in r["segments"] for w in s["words"]]
left = [(w["word"], round(w["start"], 1)) for w in words
        if w["word"].strip(" .,!?").lower() in ("um", "uh", "umm", "uhh", "er")]
print(len(left), "fillers left", left[:6])

Every Claude video editing guide

Turn a Podcast Into Shorts With Claude (2026)Turn a long podcast or video into vertical shorts with Claude Code: pick moments from the transcript, tighten, pace, reframe to 9:16 and caption. Real steps.Auto Captions and Subtitles With Claude (2026)Add captions and subtitles to video with Claude Code: word timings from Whisper, an SRT rebuilt from the cut, burned-in captions that shrink to fit, a QA gate.Auto Reframe Video to 9:16 With Claude (2026)Reframe horizontal video to vertical 9:16 with Claude Code: measure the face per shot, crop around it with ffmpeg, keep the head in frame and check every shot.Clean Up Podcast Audio and Hit -14 LUFS With ClaudeClean up podcast and talking-head audio with Claude Code and ffmpeg: high-pass, expander, denoise, EQ, 3:1 compression, de-esser, then -14 LUFS and -1 dBTP.Remove Silence and Jump Cut Video With ClaudeTrim dead air and pauses from talking-head video with Claude Code and ffmpeg: detect pauses over 0.4 s, cut them to 0.2 s and hide the jump. Commands included.Make YouTube Thumbnails With Claude (2026)Make YouTube thumbnails, Shorts covers and LinkedIn crops with Claude Code: the frame where the hook lands, the title in the middle band, four shapes.Descript Alternative: Edit With Claude Instead (2026)Replace Descript with Claude Code, ffmpeg and Whisper: filler removal, transcript cuts, captions and Studio Sound-style cleanup. Prices compared, limits stated.Opus Clip Alternative: Make Shorts With Claude (2026)Turn podcasts into shorts with Claude Code instead of Opus Clip: transcript-picked moments, 9:16 reframes, captions. Opus Clip, Riverside, Vizard prices.CapCut Alternative: Edit Talking-Head Video With ClaudeA CapCut alternative for talking-head and podcast video: Claude Code with ffmpeg and Whisper for captions, silence and filler cuts, 9:16 reframing and loudness.Submagic Alternative: Captions and Shorts With ClaudeA Submagic alternative: Claude Code writes word-timed captions, cuts silences and fillers with ffmpeg. Submagic, Captions and Vizard prices compared.

Frequently asked questions

How do I remove filler words from a video automatically?

Transcribe it with word timestamps and a prompt full of "um, uh, like" so the fillers are written down, cut each filler in the quiet gap between words, fade every join by 10 to 20 ms, and re-transcribe the result to confirm none are left. Claude Code runs all of it with ffmpeg and Whisper.

Can Claude remove ums from audio?

Yes. Claude Code reads a primed Whisper transcript of the audio, writes a keep list of every range between the fillers, and renders it with ffmpeg. It works the same on a wav, an mp3 or the audio of a video.

Why does my transcript not show um and uh?

Speech-to-text models are trained to produce clean text, so they leave disfluencies out. Passing an initial prompt that contains fillers makes Whisper transcribe them.

Will cutting fillers make the video look jumpy?

On one fixed camera, yes, unless you hide the jump. Use a camera change, a short punch-in, or a 5-frame dissolve on horizontal edits. Vertical shorts tolerate harder jump cuts.

Want finished ads, not editing sessions?

Director by Sprites is a full creative team for your startup. Every Monday you get finished, on-brand video ads, reviewed by a human. Strategy, hooks, scripts, shooting and editing, end to end. You never write a prompt. You pick the winners. Sprites ships the campaigns. In private beta.