Lip sync · mouth animation
Realistic mouth animation from one image
Realistic mouth animation comes from matching mouth shapes to the sounds being spoken: lips pressed together for M, B and P, upper teeth on the lower lip for F and V, rounded lips for OO and W, an open jaw for AH. Animators keyframe those shapes by hand. AI lip sync generates them from the audio instead, rebuilding the mouth region of one photo or character image frame by frame, so a 20 second clip is ready in about two and a half minutes on Percify.
- 1
- image of a face or a character
- 7.5 s
- of compute per second of video
- 19 s
- median clip people keep
- 69
- credits for an average 480p render

The mouth shapes that carry speech
Animators group sounds by the shape the mouth makes, not by letter. A small set of shapes covers most speech, and landing those correctly matters more than smooth motion between them. The closed shapes for M, B and P are the ones viewers notice first, because the lips have to fully meet.
| Sounds | Mouth shape |
|---|---|
| M, B, P | Lips pressed together, then released |
| F, V | Upper teeth resting on the lower lip |
| OO, W | Lips rounded and pushed forward |
| AH | Jaw dropped, mouth open wide |
| EE | Lips stretched wide, teeth close together |
| O | Lips rounded, jaw partly open |
| TH | Tongue tip between the teeth |
| L, T, D, N | Tongue touching just behind the upper teeth |
Why it is slow by hand
Hand animation means placing a shape for every syllable, then tuning the timing. Three habits separate convincing mouths from puppet mouths: animate the sounds, not the spelling; let the shape arrive a frame or two before the sound so it reads as the cause; and hold shapes instead of flapping on every letter.
That is hours of work for a minute of dialogue, and it has to be redone whenever the line changes. It is also the part of the process a model can take over completely, because the audio already contains the timing.
How AI generates the mouth instead
An AI lip sync model rebuilds the mouth region of each frame from the audio. There is no rig and no keyframe: you give it one image of a face and the sound, and the shapes come from the sound. That makes the audio the thing that decides the result, so a dub has to drive its own mouth rather than borrow the original one.
It works on stylised characters as well as photos: anime, 3D and cartoon faces all sync, and a generated character is the easy case, because nothing ever crosses its mouth. Start one in the anime avatar maker.

What breaks AI mouth animation
Across 160 source images on Percify, about 9% were run a second time, and almost always for one of these reasons. The fixes are in fix AI lip sync.
- The face is under a quarter of the frame width, so there are too few mouth pixels to rebuild.
- The head moves fast enough to compete with the mouth.
- A hand, microphone or strand of hair crosses the lips.
- A dub plays over a mouth that was generated from the original audio.
- The video has a variable frame rate, so sync drifts toward the end.
Cost and time
A 480p render averages 69 credits and a 720p render 155, on one credit meter shared with images and voices. The engine spends about 7.5 seconds of compute per second of video, and the median render takes 168 seconds. For social video, 480p on a face that fills the frame is the setting two thirds of people keep.
Animate a mouth in three steps
- 1
Start from a clear image
A photo or a character with the face large in frame and the mouth unobstructed.
- 2
Give it the audio
Write a script and pick a voice, or bring your own recording.
- 3
Check the closed shapes
Watch M, B and P. If the lips meet cleanly there, the rest of the sync is almost always right.
Mouths that match the sound
Stylised or photoreal, the mouth is generated from the audio.
Sarah
an anime look
Tyler
a 3D character
Animated character
a cartoon child, lip synced
Frequently asked questions
What are the basic mouth shapes for animation?
Lips pressed for M, B and P; upper teeth on the lower lip for F and V; rounded lips for OO and W; an open jaw for AH; stretched lips for EE; rounded and partly open for O; the tongue between the teeth for TH; and the tongue behind the upper teeth for L, T, D and N. That small set covers most speech.
How do I make mouth movements look realistic?
Match shapes to sounds rather than letters, make sure the closed shapes fully close, let each shape arrive a frame or two before its sound, and hold shapes instead of moving on every letter. Or generate the mouth from the audio with AI lip sync.
Can AI animate a mouth from a picture?
Yes. Give it one photo or character image and the audio, and it rebuilds the mouth frame by frame. On Percify that takes about 7.5 seconds of compute per second of video.
Does AI mouth animation work on cartoon and anime characters?
Yes. Generated characters are the easy case, because nothing crosses the mouth. Make one in the anime avatar maker, then give it a voice.
Why do the mouth shapes look wrong on a dubbed video?
Because the mouth was generated from the original language while the dub plays. The timing matches, the shapes do not. Generate the mouth from the translated track.
Keep reading
Give a face a voice
One image and the words. The mouth follows the sound.