Lip sync · mouth animation

Realistic mouth animation from one image

Realistic mouth animation comes from matching mouth shapes to the sounds being spoken: lips pressed together for M, B and P, upper teeth on the lower lip for F and V, rounded lips for OO and W, an open jaw for AH. Animators keyframe those shapes by hand. AI lip sync generates them from the audio instead, rebuilding the mouth region of one photo or character image frame by frame, so a 20 second clip is ready in about two and a half minutes on Percify.

Free to start · no credit card
1
image of a face or a character
7.5 s
of compute per second of video
19 s
median clip people keep
69
credits for an average 480p render
Line drawing of three stacked layers: audio at the base, a face above it, the finished video on top

The mouth shapes that carry speech

Animators group sounds by the shape the mouth makes, not by letter. A small set of shapes covers most speech, and landing those correctly matters more than smooth motion between them. The closed shapes for M, B and P are the ones viewers notice first, because the lips have to fully meet.

SoundsMouth shape
M, B, PLips pressed together, then released
F, VUpper teeth resting on the lower lip
OO, WLips rounded and pushed forward
AHJaw dropped, mouth open wide
EELips stretched wide, teeth close together
OLips rounded, jaw partly open
THTongue tip between the teeth
L, T, D, NTongue touching just behind the upper teeth

Why it is slow by hand

Hand animation means placing a shape for every syllable, then tuning the timing. Three habits separate convincing mouths from puppet mouths: animate the sounds, not the spelling; let the shape arrive a frame or two before the sound so it reads as the cause; and hold shapes instead of flapping on every letter.

That is hours of work for a minute of dialogue, and it has to be redone whenever the line changes. It is also the part of the process a model can take over completely, because the audio already contains the timing.

How AI generates the mouth instead

An AI lip sync model rebuilds the mouth region of each frame from the audio. There is no rig and no keyframe: you give it one image of a face and the sound, and the shapes come from the sound. That makes the audio the thing that decides the result, so a dub has to drive its own mouth rather than borrow the original one.

It works on stylised characters as well as photos: anime, 3D and cartoon faces all sync, and a generated character is the easy case, because nothing ever crosses its mouth. Start one in the anime avatar maker.

Line drawing of a portrait photo feeding into a waveform, the audio that drives the mouth

What breaks AI mouth animation

Across 160 source images on Percify, about 9% were run a second time, and almost always for one of these reasons. The fixes are in fix AI lip sync.

  • The face is under a quarter of the frame width, so there are too few mouth pixels to rebuild.
  • The head moves fast enough to compete with the mouth.
  • A hand, microphone or strand of hair crosses the lips.
  • A dub plays over a mouth that was generated from the original audio.
  • The video has a variable frame rate, so sync drifts toward the end.

Cost and time

A 480p render averages 69 credits and a 720p render 155, on one credit meter shared with images and voices. The engine spends about 7.5 seconds of compute per second of video, and the median render takes 168 seconds. For social video, 480p on a face that fills the frame is the setting two thirds of people keep.

Animate a mouth in three steps

  1. 1

    Start from a clear image

    A photo or a character with the face large in frame and the mouth unobstructed.

  2. 2

    Give it the audio

    Write a script and pick a voice, or bring your own recording.

  3. 3

    Check the closed shapes

    Watch M, B and P. If the lips meet cleanly there, the rest of the sync is almost always right.

Mouths that match the sound

Stylised or photoreal, the mouth is generated from the audio.

Frequently asked questions

What are the basic mouth shapes for animation?

Lips pressed for M, B and P; upper teeth on the lower lip for F and V; rounded lips for OO and W; an open jaw for AH; stretched lips for EE; rounded and partly open for O; the tongue between the teeth for TH; and the tongue behind the upper teeth for L, T, D and N. That small set covers most speech.

How do I make mouth movements look realistic?

Match shapes to sounds rather than letters, make sure the closed shapes fully close, let each shape arrive a frame or two before its sound, and hold shapes instead of moving on every letter. Or generate the mouth from the audio with AI lip sync.

Can AI animate a mouth from a picture?

Yes. Give it one photo or character image and the audio, and it rebuilds the mouth frame by frame. On Percify that takes about 7.5 seconds of compute per second of video.

Does AI mouth animation work on cartoon and anime characters?

Yes. Generated characters are the easy case, because nothing crosses the mouth. Make one in the anime avatar maker, then give it a voice.

Why do the mouth shapes look wrong on a dubbed video?

Because the mouth was generated from the original language while the dub plays. The timing matches, the shapes do not. Generate the mouth from the translated track.

Keep reading

Give a face a voice

One image and the words. The mouth follows the sound.

Free to start · no credit card