AI Dubbing · 2026 Guide
AI Video Dubbing: The Step-by-Step Guide
To dub a video with AI, upload it, choose a target language, and generate — the model transcribes the speech, translates it, synthesises new audio and re-syncs the speaker’s lips to match. On Percify this runs as a single generation across 178 languages and dialects, with a voice-preserving option that keeps the original speaker’s voice in the new language.
Below is the full process, what actually happens at each stage, how the dubbing models differ, and what to check before you localise a campaign.
How to dub a video with AI, step by step
- Upload your video. Upload an MP4 with clear speech. One speaker and minimal background music gives the most reliable dub.
- Choose the target language. Pick from 178 output languages and dialects. Choose the specific regional variant if you have one, since phrasing differs between them.
- Pick stock voice or voice clone. Use a synthesised voice, or a voice-preserving model that keeps the original speaker’s timbre in the new language.
- Generate the dub. The model transcribes, translates, synthesises the new audio, time-aligns it to the original, and regenerates the speaker’s lip movements to match.
- Review and download. Check the first and last few seconds for drift, then download. Paid plans include commercial usage rights.
What happens inside an AI dub
Dubbing is five distinct operations that older workflows performed separately, each with its own tool and hand-off:
- Speech recognition — the original audio is transcribed with timings.
- Translation — the transcript is translated, keeping the delivery natural rather than literal.
- Voice synthesis — a stock voice or a clone of the original speaker performs the new script.
- Time alignment — the new audio is stretched or compressed to the original timing, since translations rarely match the source length.
- Lip-sync — the mouth is regenerated to match the new phonemes. This is the step that separates a dub from a voice-over.
AI dubbing and translation models compared
| Model | Best for | Languages | Cost |
|---|---|---|---|
| HeyGen Video Translate | Best for full video dubbing | 178 languages and dialects | 180 credits (90s minimum) |
| ElevenLabs Dubbing | Best for voice-preserving dubs | 18 languages | 72 credits per minute |
| AI Image Translator | Best for on-image text | 34 languages | 12 credits |
| Qwen Image Translate | Cheapest image translation | 9 languages | 2 credits |
HeyGen Video Translate: Translates a talking video into another language and re-syncs the lips to the new audio, so the result looks natively filmed.
ElevenLabs Dubbing: Keeps the original speaker’s voice character in the target language, which matters when the presenter is your brand.
AI Image Translator: Replaces text inside an image while preserving font, spacing and layout — for thumbnails, ad creative and packaging.
Qwen Image Translate: OCR-based in-image translation with domain hints and custom terminology lists for consistent product naming.
AI dubbing tools for marketing teams
For marketing, dubbing is a distribution multiplier: one shoot becomes a campaign in every market you sell in. Three things decide whether that works in practice. Lip-sync quality, because an unsynced dub reads as spam in a feed. Commercial licensing, which Percify includes on paid plans. And bulk throughput — localising twelve markets by hand is a week of work, so teams run it through the Percify API instead. Remember the creative around the video too: thumbnails, captions and ad images carry text that stays in the source language unless you translate it, which is what the image models above are for.
Dubbing vs subtitles vs voice-over
Subtitles leave the audio alone and add translated text — cheapest, but most social viewers will not read them. Voice-over replaces the audio while the lips keep moving to the original language, which viewers notice immediately. Lip-sync dubbing replaces the audio and regenerates the mouth, so the speaker appears to have filmed in that language. For paid ads and creator content, dubbing is the only one of the three that does not visibly signal a translation. If you need the speaker to stay recognisable, pair it with voice cloning.
How to get a clean dub
- Use source audio with one speaker, little background music and no overlapping dialogue.
- Keep the face visible and reasonably front-facing — lip-sync degrades on sharp profile angles.
- Pick the regional variant, not just the language, where the phrasing differs.
- Check the first and last seconds, where time-alignment drift shows up first.
- Translate the on-screen text and thumbnail too, or the video arrives in a half-localised frame.
Common pain points in AI video dubbing, and how to solve them
Most dubbing guides describe the happy path. These are the six failures that actually come up, what causes each one, and what to do about it.
The lips drift out of sync toward the end of the clip
Why it happens. Translated speech is rarely the same length as the original. German and Spanish run longer than English, Japanese runs shorter, so the new audio has to be stretched or compressed to fit the original timing. Small per-sentence corrections accumulate, which is why drift shows up at the end rather than the start.
The fix. Check the last few seconds first, not the first few. If it has drifted, cut the source into segments at natural pauses and dub each one, so timing resets at every boundary instead of accumulating across the whole video.
The dubbed voice sounds nothing like the original speaker
Why it happens. Standard dubbing synthesises a stock voice — it translates what was said, not who said it. Voice preservation is a different model, and it covers 18 languages rather than the full 178.
The fix. Use the voice-preserving model where your target language is among its 18. Outside that set the timbre cannot be kept, so pick one stock voice and reuse it across every language to stay consistent.
Multiple speakers get blended into one voice
Why it happens. Dubbing quality depends on how cleanly speakers are separated in the source audio. Overlapping dialogue and background music leave the model no clean boundary to work from.
The fix. Dub each speaker’s segments separately and reassemble them. If you control the shoot, record dialogue on separate tracks and keep music off the speech track.
Product names and jargon come back mistranslated
Why it happens. Translation treats a brand or product name as an ordinary word and localises it, which is usually wrong for names and technical terms.
The fix. Say the product name in the source audio exactly as it should appear in every language, and review the first dub for terminology before running the rest of the batch.
The video is dubbed but the on-screen text is still in English
Why it happens. Dubbing replaces spoken audio. Text baked into the footage, thumbnails and ad creative is a separate job.
The fix. Run image translation on those assets. It redraws the text in the original font and layout rather than pasting a caption over it, in 34 languages, with a cheaper 9-language model for bulk work.
A short clip cost more than expected
Why it happens. Full video dubbing bills a 90-second minimum at 180 credits, so a 20-second clip and a 90-second clip cost the same.
The fix. Batch short clips into one video before dubbing, then split the result — or use voice-preserving dubbing, which bills per minute at 72 credits.
Frequently asked questions
How do you dub a video with AI, step by step?
Upload your video, choose the target language, and generate. On Percify: open the HeyGen Video Translate model, upload an MP4 with clear speech, pick one of 178 output languages, and run it. The model transcribes the speech, translates it, generates the dubbed audio and re-syncs the speaker’s lips to match. A 90-second clip costs 180 credits. You then download the finished video — no editing timeline and no separate subtitle step.
What are the steps in AI video dubbing?
AI video dubbing has five steps: (1) speech recognition transcribes the original audio, (2) the transcript is translated into the target language, (3) a voice is synthesised — either a stock voice or a clone of the original speaker, (4) the new audio is time-aligned to the original timing, and (5) the speaker’s lip movements are regenerated to match the new audio. Percify runs all five in a single generation.
What is the difference between AI dubbing and AI subtitles?
Subtitles add translated text on screen and leave the original audio untouched. Dubbing replaces the spoken audio with the target language, and lip-sync dubbing also regenerates the mouth movements so the speaker appears to be speaking that language. Dubbing performs far better for social and ad content, where most viewers will not read subtitles.
Can AI dubbing keep the original speaker’s voice?
Yes. Voice-preserving dubbing clones the timbre of the original speaker and speaks the translated script in that voice. On Percify the ElevenLabs Dubbing model does this across 18 languages, which keeps a founder or presenter recognisable across every localised version.
What are the best AI dubbing tools for marketing teams?
Marketing teams should choose a dubbing tool on three criteria: lip-sync quality, language coverage, and whether output is commercially licensed. Percify covers all three — 178 languages for full video dubbing, voice-preserving dubs in 18 languages, in-image text translation for thumbnails and ad creative, commercial usage rights on paid plans, and an API so one campaign can be localised in bulk.
How much does AI video dubbing cost?
Percify prices dubbing in credits rather than per seat. Full video dubbing with lip-sync costs 180 credits for a 90-second minimum, voice-preserving dubbing costs 72 credits per minute, and in-image text translation costs 2–12 credits per image. You can try the models before subscribing.
Can I translate the text inside an image or thumbnail?
Yes, and it is a separate job from video dubbing. Image translation extracts the text from an image, translates it, and redraws it in the original font, spacing and layout so the creative still looks designed rather than patched. Percify offers this in 34 languages, and a cheaper 9-language model for high-volume work.
What are the common pain points in video translation and dubbing?
The six that come up most: lip-sync drift late in the clip caused by translated speech running a different length than the original; a dubbed voice that no longer sounds like the speaker, because voice preservation covers 18 languages while standard dubbing covers 178; multiple speakers blending together when the source audio is not cleanly separated; product names and jargon being localised when they should be left alone; on-screen text staying in the original language because dubbing only replaces audio; and short clips costing more than expected because full dubbing bills a 90-second minimum.
Why does AI dubbing lip-sync drift out of time?
Because translated speech is rarely the same length as the source. German and Spanish run longer than English and Japanese runs shorter, so the generated audio is stretched or compressed to fit the original timing. Those per-sentence corrections accumulate, which is why drift is visible at the end of a clip rather than the beginning. Cutting the video into segments at natural pauses and dubbing each separately resets the timing at every boundary.
How do you fix bad voice quality or a wrong accent in an AI dub?
Choose the model deliberately. Standard dubbing synthesises a stock voice, so it will not match the original speaker in either timbre or accent. Voice-preserving dubbing clones the original speaker across 18 languages. Outside those languages the timbre cannot be preserved, so the consistent choice is one stock voice reused across every language rather than a different one per market.
Does AI dubbing work for videos with multiple speakers?
Quality depends on how cleanly the speakers are separated in the source audio. Dubbing works best on a single speaker with clear, unmixed dialogue and little background music. For multi-speaker footage, dub each speaker’s segments separately and reassemble them for the most reliable result.
Dub your first video free
Translate and lip-sync a video into 178 languages. No credit card to try, commercial rights on paid plans.
Try AI dubbing free