Skip to content
Back to blog
editingfiller-wordstranscript

How to automatically cut filler words and silences from a video

Learn how to automatically cut filler words, ums, and dead-air silences from a long-form video by editing the transcript instead of scrubbing the timeline.

By Franck, founder of B-Rollss · July 23, 2026 · 4 min read

Record yourself talking for forty minutes and play it back. You will hear it immediately: "um," "so," "you know," "like," the three-second pause while you remembered the next point, the sentence you started twice. None of it is a content problem. It is a cleanup problem — and for long-form creators it is the single most time-consuming, least creative part of the whole edit.

Here is how to cut filler words and silences from a video without spending your evening dragging clips around a timeline.

Why the timeline is the slow way

The traditional method is brutal: play the video, hear an "um," stop, find the exact in and out points, cut, close the gap, keep going. For a fifteen-minute finished video pulled from forty minutes of raw footage, you are doing that hundreds of times. It is precise but mind-numbing, and one mis-placed cut leaves an audible clip or a jump that you only catch on the third review.

The problem is that you are working in the wrong medium. Audio filler is a *language* problem — it lives in the words. So the fastest place to fix it is the words, not the waveform.

Edit the transcript, not the video

The shift that makes this fast is transcript-based editing. You transcribe the recording down to the word, with each word tied to its exact moment in the video. Now the edit becomes text editing:

  • Delete "um" in the transcript and the exact slice of video attached to it disappears.
  • Highlight a rambling sentence, delete it, and the video closes the gap cleanly.
  • Read the transcript top to bottom — if it reads tight, it watches tight.

This flips the whole task. Instead of hunting through a waveform for sounds, you scan text for words, which your brain does far faster. You catch the "you know" you would have missed by ear, and you never fumble an in-point again because there is no in-point to fumble.

The three things worth cutting

Not everything should go. Over-cutting makes a video feel robotic and breathless. In practice there are three targets:

  1. Filler words — "um," "uh," "like," "so," "basically," "you know." These add nothing and cutting them tightens pacing instantly.
  2. Long silences and dead air — the gaps where you thought, checked notes, or drifted. Trimming the long ones keeps momentum; leaving short natural pauses keeps you human.
  3. Botched takes and false starts — the sentence you began, abandoned, and restarted. Keep the clean take, drop the rest.

What you should *not* do is remove every breath and micro-pause. Speech needs rhythm. The goal is a cut that feels tight and natural, not one that sounds like a machine gun.

Doing it in one pass instead of one word at a time

Even in a transcript, cutting hundreds of fillers by hand is a chore. This is where automation earns its place: a pass that scans the whole transcript, flags every filler and every over-long silence, and removes them in one go — while you keep the ability to review and undo anything it touched.

That combination matters. Fully manual is too slow; fully automatic with no review is too risky, because sometimes an "um" is doing rhythmic work or a pause is intentional. The right tool does the tedious bulk removal automatically and still lets you overrule it word by word.

How I do it now

I will be direct about what I built, since this is the exact problem it solves for my own videos.

With Brollss, I upload the recording and it transcribes the whole thing word by word. From that editable transcript I clean up mechanically: a whole-transcript pass tightens filler and dead air in one shot, and if I want finer control I cut a word in the text and its video slice is gone. For anything fuzzier — "tighten the intro," "this section drags" — an AI co-editor takes the instruction in plain language and applies it as precise, reviewable timeline edits.

The point is not that software decides what your video says. You still own every call. The point is that the most mechanical hour of the edit — the ums, the silences, the false starts — stops being manual labor and becomes a quick review.

And cleaning the audio is only half the evening. The other half is the visual layer: charts, code, quotes, the things that keep a talking head watchable. Brollss generates that from the same transcript too — but the filler-and-silence cleanup alone is the part most creators feel first, because it is the part they hate most.

Try it on a real recording

If the boring cleanup pass is what stalls your uploads, this is exactly why I built Brollss. We are at founding-access stage — I am onboarding early creators one at a time and reading all the feedback myself.

Request founding access, bring one messy forty-minute recording, and see how tight it reads once the filler is gone.

Ready to edit faster?