Text-Based Video Editor: edit your video like a document
Drop a talking-head video, get a word-by-word transcript, and delete words to cut them from the video. Remove fillers, retakes and long pauses in one click, add captions that highlight each word, and export an MP4. Everything runs on your device — no upload, no watermark, no minute limit.
Free · No sign-upHow to use Text-Based Video Editor
- 1. Drop or choose a video (MP4, MOV, WebM…) or an audio file. It opens on your device — nothing is uploaded.
- 2. Press Transcribe. The first time, the speech model downloads once (the size is shown, about 77 MB for Base) and is cached. You get a transcript where every word is timed — click any word to jump there.
- 3. Edit the text: select words and press Delete to cut them (press Delete again to restore, Ctrl/⌘+Z to undo). Or use Remove fillers, Remove retakes and Tighten silences — each shows what it found so you can review it first.
- 4. Pick a caption style, download SRT / VTT or burn the captions in, then press Export video to get an MP4 with no watermark. Signed in? Send it straight to the Scheduler.
Frequently asked questions
Is my video uploaded?
No. The video is decoded, transcribed, cut and encoded inside your browser. The only downloads are the speech model (41–564 MB depending on the model you pick, once, then cached) and the video libraries. If you choose Send to Scheduler, the finished video is uploaded to your own Scheduler media library — only when you press that button.
How does cutting by text work?
Whisper writes down every word with its start and end time. When you delete words, the editor cuts the video from the middle of the pause before them to the middle of the pause after them, moves each cut edge to the quietest moment nearby, and adds a short audio crossfade (30–60 ms) so there are no clicks.
Which filler words are removed?
Hesitation sounds like um, uh, erm and hmm are always suggested. Words that can be real words, like like, you know and I mean, are only suggested when there is a pause on both sides. Nothing is cut until you review the list and press Cut, and Undo brings anything back.
How are retakes found?
When you start a sentence, stop and say it again, the earlier attempt is suggested for removal and the later one is kept: a sentence that is the start of the next one, two neighbouring sentences with mostly the same words, or the same three or more words said twice in a row.
Is there a watermark or a time limit?
No watermark and no minute cap. The only limit is your device's memory: on a computer, videos up to about an hour and a half and 2 GB usually work; on phones, shorter clips. You see a warning before you start if a file looks too big for this device.
How fast is it?
On a laptop processor, the Base model transcribes about 7 to 9 times faster than real time, so a 10-minute video takes about 1 to 1.5 minutes; Precise word timing adds roughly 3 times that. Rendering uses the hardware video encoder through WebCodecs when your browser has it, otherwise ffmpeg, which is slower.