AI Video Glossary

The AI video terms that actually come up, defined in plain language, and for each one, the mistake it usually causes.

Read Blog

Avatars

AI avatar

A synthetic on-screen presenter generated from a photo or a text description, animated to speak audio you supply. Unlike a stock video clip, an avatar can say anything you write without a reshoot.

Where it goes wrong. An avatar built from a single low-resolution or heavily filtered photo inherits those flaws in every video it appears in. The source image sets the ceiling on quality.

Talking head

A shot framed on a single speaker from the shoulders up, addressing the camera directly. It is the format most AI avatar tools are optimised for, because the face carries the entire performance.

Where it goes wrong. Wider framing gives a model more body to invent and more room to get it wrong. Hands and shoulders are where unnatural motion shows first.

Digital twin

An avatar built to resemble a specific real person, usually combining their likeness with a clone of their voice, so one presenter can appear in far more content than they could film.

Where it goes wrong. A likeness of a real person needs that person’s consent. Building one of someone else, a competitor, a celebrity, a colleague, is the line between a digital twin and a deepfake.

Motion

Lip sync

Regenerating a speaker’s mouth movements so they match a given audio track. Good lip sync drives the whole face from the audio, jaw, cheeks and brow, rather than animating the lip region alone.

Where it goes wrong. A frozen face around a moving mouth reads as artificial within seconds, even when the lip timing itself is perfectly correct.

Phoneme

The smallest unit of sound in speech. Lip-sync models segment audio into phonemes to decide what shape the mouth should make at each moment.

Where it goes wrong. Background music and overlapping voices blur the boundaries between phonemes, so noisy source audio produces mistimed mouth shapes no matter which engine you use.

Viseme

The visual counterpart of a phoneme: the mouth shape that corresponds to a sound. Several phonemes share one viseme, which is why lip reading is ambiguous and why lip sync can look right without being phonetically exact.

Temporal drift

The gradual loss of alignment between generated video and its audio across a long take. Small per-sentence timing corrections accumulate, so the opening of a clip looks correct while the ending does not.

Where it goes wrong. Review the end of a generated clip first, not the beginning. Splitting a long take at natural pauses resets timing at every boundary instead of letting the error compound.

Voice

Voice cloning

Building a synthetic voice that reproduces the timbre of a specific speaker from a sample of their speech, so new scripts can be spoken in that voice.

Where it goes wrong. Cloning is only as good as the reference audio. A sample recorded in a noisy room bakes that room into every line the clone ever speaks.

Text to speech (TTS)

Generating spoken audio from written text using a synthetic voice. It differs from voice cloning in that the voice is a stock one rather than a copy of a particular person.

Voice preservation

Dubbing that keeps the original speaker’s timbre in the target language, so a founder or presenter stays recognisable across every localised version rather than being replaced by a stock voice.

Where it goes wrong. Voice preservation covers fewer languages than standard dubbing, 18 against 178 on Percify. Outside that set the timbre cannot be kept, so the consistent choice is one stock voice reused across all markets.

Localisation

AI dubbing

Replacing a video’s spoken audio with a translated version, and regenerating the speaker’s lip movements to match. Five stages run in sequence: transcription, translation, voice synthesis, time alignment, and lip-sync.

Where it goes wrong. Translated speech is rarely the same length as the original, German and Spanish run longer than English, Japanese shorter, so the audio is stretched or compressed to fit, which is the root cause of most sync problems.

Subtitles vs dubbing

Subtitles add translated text on screen and leave the original audio intact. Dubbing replaces the audio itself, and lip-sync dubbing also regenerates the mouth so the speaker appears to have filmed in that language.

Where it goes wrong. Subtitles depend on the viewer reading them, which most social and ad audiences will not do. Dubbing is the only option that does not visibly signal a translation.

Image translation

Extracting text baked into an image, translating it, and redrawing it in the original font, spacing and layout, used for thumbnails and ad creative, where a pasted caption would look patched rather than designed.

Where it goes wrong. Dubbing replaces spoken audio only. On-screen text, thumbnails and creative are a separate job, and skipping them ships a video in a half-localised frame.

Generation

Text to video

Generating a video clip directly from a written description, with no source footage or photo. The model invents the subject, the motion and the scene together.

Where it goes wrong. Because nothing anchors the output, the same prompt run twice returns different people and places. Text to video is weak wherever a subject has to stay consistent between shots.

Image to video

Animating a still image into a video clip, keeping the subject in the photo and generating motion around it. This is the mode most avatar work uses, because the person stays fixed.

Diffusion model

A generative model that starts from random noise and removes it step by step until an image or video emerges. Most current image and video generators work this way.

Where it goes wrong. More steps generally means better quality and longer generation time. The trade-off is a cost decision rather than a quality bug.

Prompt

The written instruction given to a generative model. For video it usually describes the subject, the framing, the motion and the lighting.

Where it goes wrong. Prompts describing a style rather than a scene tend to return the source image barely changed. Name the shot, not just the aesthetic.

Inference

A single run of a trained model to produce an output. What you are billed for when you generate a video is inference time on the provider’s hardware.

Commercial

Credits

A usage unit that prices generation by what it actually costs to run, rather than per seat. Different models consume different amounts: full video dubbing costs 180 credits for a 90-second minimum, voice-preserving dubbing 72 credits per minute, and image translation 2 to 12 credits per image.

Where it goes wrong. A minimum-duration charge means a 20-second clip and a 90-second clip can cost the same. Batching short clips into one job before dubbing avoids paying the minimum repeatedly.

Watermark

A visible mark burned into generated output identifying the tool that made it. Free tiers commonly apply one; paid plans generally remove it.

Where it goes wrong. A watermark is not the same question as usage rights. A video can be watermark-free and still not be licensed for commercial use, check both before running it as an ad.

Commercial usage rights

Permission to use generated output in work that makes money, ads, client deliverables, monetised channels. On Percify these are included on every plan.

Where it goes wrong. Rights to the generated output do not extend to the likeness inside it. Using a real person’s face or voice still needs their consent, whatever the plan allows.

Ethics

Deepfake

Synthetic media depicting a real person saying or doing something they did not, made without their consent. The technique overlaps with avatar generation; consent is what separates them.

Where it goes wrong. The distinction is consent and disclosure, not the tool. The same model produces a legitimate digital twin or a deepfake depending entirely on whose likeness it is and whether they agreed.

Capability figures on this page, 178 dubbing languages, 18 for voice preservation, 34 for image translation, and the credit costs, match the dubbing guide and the model catalog.