AI video dubbing: add an English voice-over track
Turn your video's subtitles into an English dub read by a stock AI voice — translated line by line, voiced on your device, fitted to the original timing and mixed over the original music and ambience. Download an MP4 with the dub track, the dub audio for YouTube's multi-language audio, and English subtitles. No upload, no watermark, and never a copy of anyone's voice.
Free · No sign-upHow to dub a video into English with AI
- 1. Drop your video (or audio) and its SRT or VTT subtitles. No subtitles? Press Transcribe — Whisper writes them on your device (the speech model downloads once, 41–250 MB).
- 2. Pick the language of the subtitles and press Translate into English. Each line is translated on its own, so the timing stays. Edit any line — shorter lines are easier to fit.
- 3. Choose a stock AI voice (US or UK English), the pace and how much faster long lines may be spoken, then press Make the English dub. The voice model (~92 MB) downloads once and is cached.
- 4. Set how much the original sound is lowered under the voice, preview it with the picture, and download the MP4, the dub audio track (WAV, for YouTube's multi-language audio) or English subtitles (SRT).
How the dub stays in sync with the picture
A dub has to fit the time the original speaker used. Every English line is voiced separately, its silence at both ends is trimmed, and it is placed where its subtitle starts. When a line is longer than the time before the next one, the tool first lets it start up to 0.3 seconds early, then asks the voice model to speak it faster — up to the limit you choose, 1.25× by default — so the voice stays natural instead of sounding stretched. If a line still doesn't fit, its end is faded out (or, if you choose, it may run up to a second late and the next line waits). The timing report lists every line that didn't fit, so you can shorten the English and update just those lines.
The original soundtrack isn't thrown away. While the AI voice speaks, the original is lowered by 15 dB (adjustable, or muted) with smooth fades, and it comes back up in the pauses, so music, laughter and room sound keep the video alive. The original voice is lowered rather than removed. Translations into English usually come out shorter than Spanish, French or German, so most lines fit at normal pace; very fast talkers need a higher speed limit or a few shortened lines.
Being open about synthetic audio keeps viewers' trust and follows platform rules: disclose it with YouTube's altered or synthetic content setting, TikTok's AI-generated label or Meta's AI info, and say “AI-dubbed” in the description. Need subtitles in many languages instead of a voice? Use the Multilingual Subtitle Pack. To read any text aloud, try Text to Speech; to cut a talking-head video by editing its transcript, the Text-Based Video Editor.
Frequently asked questions
Is this voice cloning?
No. The dub is read by one of 13 stock voices that come with the Kokoro voice model (US and UK English, female and male). The tool never records, copies or imitates the voice of the person in your video, and there is no option to do so.
Which languages can it dub into?
English, with US or UK voices. Your subtitles can be in any of 35 languages; each line is translated into English first. The voice model also has Spanish, French, Italian, Portuguese, Hindi, Japanese and Chinese voices, but in a browser those need a pronunciation engine that isn't available under an open licence yet, and it has no German voice. For other languages, make translated subtitles with the Multilingual Subtitle Pack.
How does it keep the dub in sync?
Each English line is voiced, its silence trimmed, and placed where its subtitle starts. A line that is too long may start up to 0.3 s early, then is spoken faster (up to the limit you set, 1.25× by default, re-voiced by the model at that speed so it still sounds natural). If it still doesn't fit, its end is faded out, or, if you prefer, it runs up to 1 s late. Every line that doesn't fit is listed so you can shorten it.
What happens to the music and background sound?
It stays. While the AI voice speaks, the original sound is lowered (15 dB by default) and it comes back in the pauses, so music and ambience keep playing. The original voice is lowered, not removed. You can also mute the original completely under the voice.
Do I have to label the video as AI?
Tell viewers the voice is synthetic. YouTube asks creators to disclose realistic altered or synthetic content, TikTok has an AI-generated content label and Meta shows AI info; in the EU, the AI Act (Article 50) sets transparency rules for realistic AI-generated audio. The MP4's audio track is named “English (AI voice)” and a comment in the file says it was made with a synthetic voice.
Is my video uploaded, and is there a limit?
No upload: transcription, the voice model and the mixing run in your browser. Only the subtitle text goes to a translation service when Chrome's on-device translator isn't available. The voice model (~92 MB) downloads once and is cached. No watermark and no length cap; very long videos are limited by your device's memory, and the tool tells you when a file is likely too long.