AI Dubbing · 2026 Guide

AI Video Dubbing: The Step-by-Step Guide

To dub a video with AI, upload it, choose a target language, and generate — the model transcribes the speech, translates it, synthesises new audio and re-syncs the speaker’s lips to match. On Percify this runs as a single generation across 178 languages and dialects, with a voice-preserving option that keeps the original speaker’s voice in the new language.

Below is the full process, what actually happens at each stage, how the dubbing models differ, and what to check before you localise a campaign.

How to dub a video with AI, step by step

  1. Upload your video. Upload an MP4 with clear speech. One speaker and minimal background music gives the most reliable dub.
  2. Choose the target language. Pick from 178 output languages and dialects. Choose the specific regional variant if you have one, since phrasing differs between them.
  3. Pick stock voice or voice clone. Use a synthesised voice, or a voice-preserving model that keeps the original speaker’s timbre in the new language.
  4. Generate the dub. The model transcribes, translates, synthesises the new audio, time-aligns it to the original, and regenerates the speaker’s lip movements to match.
  5. Review and download. Check the first and last few seconds for drift, then download. Paid plans include commercial usage rights.

What happens inside an AI dub

Dubbing is five distinct operations that older workflows performed separately, each with its own tool and hand-off:

  1. Speech recognition — the original audio is transcribed with timings.
  2. Translation — the transcript is translated, keeping the delivery natural rather than literal.
  3. Voice synthesis — a stock voice or a clone of the original speaker performs the new script.
  4. Time alignment — the new audio is stretched or compressed to the original timing, since translations rarely match the source length.
  5. Lip-sync — the mouth is regenerated to match the new phonemes. This is the step that separates a dub from a voice-over.

AI dubbing and translation models compared

ModelBest forLanguagesCost
HeyGen Video TranslateBest for full video dubbing178 languages and dialects180 credits (90s minimum)
ElevenLabs DubbingBest for voice-preserving dubs18 languages72 credits per minute
AI Image TranslatorBest for on-image text34 languages12 credits
Qwen Image TranslateCheapest image translation9 languages2 credits

HeyGen Video Translate: Translates a talking video into another language and re-syncs the lips to the new audio, so the result looks natively filmed.

ElevenLabs Dubbing: Keeps the original speaker’s voice character in the target language, which matters when the presenter is your brand.

AI Image Translator: Replaces text inside an image while preserving font, spacing and layout — for thumbnails, ad creative and packaging.

Qwen Image Translate: OCR-based in-image translation with domain hints and custom terminology lists for consistent product naming.

AI dubbing tools for marketing teams

For marketing, dubbing is a distribution multiplier: one shoot becomes a campaign in every market you sell in. Three things decide whether that works in practice. Lip-sync quality, because an unsynced dub reads as spam in a feed. Commercial licensing, which Percify includes on paid plans. And bulk throughput — localising twelve markets by hand is a week of work, so teams run it through the Percify API instead. Remember the creative around the video too: thumbnails, captions and ad images carry text that stays in the source language unless you translate it, which is what the image models above are for.

Dubbing vs subtitles vs voice-over

Subtitles leave the audio alone and add translated text — cheapest, but most social viewers will not read them. Voice-over replaces the audio while the lips keep moving to the original language, which viewers notice immediately. Lip-sync dubbing replaces the audio and regenerates the mouth, so the speaker appears to have filmed in that language. For paid ads and creator content, dubbing is the only one of the three that does not visibly signal a translation. If you need the speaker to stay recognisable, pair it with voice cloning.

How to get a clean dub

Frequently asked questions

How do you dub a video with AI, step by step?

Upload your video, choose the target language, and generate. On Percify: open the HeyGen Video Translate model, upload an MP4 with clear speech, pick one of 178 output languages, and run it. The model transcribes the speech, translates it, generates the dubbed audio and re-syncs the speaker’s lips to match. A 90-second clip costs 180 credits. You then download the finished video — no editing timeline and no separate subtitle step.

What are the steps in AI video dubbing?

AI video dubbing has five steps: (1) speech recognition transcribes the original audio, (2) the transcript is translated into the target language, (3) a voice is synthesised — either a stock voice or a clone of the original speaker, (4) the new audio is time-aligned to the original timing, and (5) the speaker’s lip movements are regenerated to match the new audio. Percify runs all five in a single generation.

What is the difference between AI dubbing and AI subtitles?

Subtitles add translated text on screen and leave the original audio untouched. Dubbing replaces the spoken audio with the target language, and lip-sync dubbing also regenerates the mouth movements so the speaker appears to be speaking that language. Dubbing performs far better for social and ad content, where most viewers will not read subtitles.

Can AI dubbing keep the original speaker’s voice?

Yes. Voice-preserving dubbing clones the timbre of the original speaker and speaks the translated script in that voice. On Percify the ElevenLabs Dubbing model does this across 18 languages, which keeps a founder or presenter recognisable across every localised version.

What are the best AI dubbing tools for marketing teams?

Marketing teams should choose a dubbing tool on three criteria: lip-sync quality, language coverage, and whether output is commercially licensed. Percify covers all three — 178 languages for full video dubbing, voice-preserving dubs in 18 languages, in-image text translation for thumbnails and ad creative, commercial usage rights on paid plans, and an API so one campaign can be localised in bulk.

How much does AI video dubbing cost?

Percify prices dubbing in credits rather than per seat. Full video dubbing with lip-sync costs 180 credits for a 90-second minimum, voice-preserving dubbing costs 72 credits per minute, and in-image text translation costs 2–12 credits per image. You can try the models before subscribing.

Can I translate the text inside an image or thumbnail?

Yes, and it is a separate job from video dubbing. Image translation extracts the text from an image, translates it, and redraws it in the original font, spacing and layout so the creative still looks designed rather than patched. Percify offers this in 34 languages, and a cheaper 9-language model for high-volume work.

Does AI dubbing work for videos with multiple speakers?

Quality depends on how cleanly the speakers are separated in the source audio. Dubbing works best on a single speaker with clear, unmixed dialogue and little background music. For multi-speaker footage, dub each speaker’s segments separately and reassemble them for the most reliable result.

Dub your first video free

Translate and lip-sync a video into 178 languages. No credit card to try, commercial rights on paid plans.

Try AI dubbing free

More Percify guides