Your audio, spoken by a face

Upload an audio file and a photo, and get an MP4 of that face speaking it with lip sync. 2 credits a second, no watermark on any plan, and you own the result.

Turn my audio into video

Three steps

Step 1

Your audio

Upload a voiceover, a podcast clip or a voice note from your phone. Common audio files such as MP3, WAV and M4A work. No audio yet? Type a script and pick a voice instead.

Step 2

A face

Upload one clear, front facing photo: you, a presenter you have permission to use, or a character you made. One person, mouth visible, works best.

Step 3

The video

The mouth, jaw and head move with the audio, and you download an MP4. A 10 second clip is 20 credits on the fast model.

A face, not a waveform

Most audio to video converters put a still picture or a moving waveform behind your MP3 so a video platform will accept it. That is fine for a song upload, and free editors like CapCut or VEED do it well. When the video should look like someone talking, a voiceover for an ad, a clip from a podcast, a lesson, Percify lip syncs a photo to your audio so the mouth, jaw and head move with every word.

Bring your own audio, or type it

Already have a recording? Upload it as it is. Have only words? Pick one of the voices or clone your own from a short clean sample, and the photo speaks your script. Every voice model and its price is on the AI voice generator page.

What it costs

You pay per second of video. On the fast lip sync model a 15 second clip is 30 credits and a 60 second clip is 120, at 480p. The higher quality model renders 720p at 6 credits a second. The playground quotes the exact cost before anything runs, and a run that fails is refunded.

Want the photo itself to move, with no speech? That is image to video. Want a full avatar of yourself, face and voice saved once? Start at the AI avatar generator.

Give your audio a face

From 2 credits a second, no watermark, and failed runs refunded.

Turn my audio into video

Frequently asked questions

Everything you need to know before you start.

Upload your audio file and one photo of a face, and Percify lip syncs the face to the audio and returns an MP4 of that person speaking it. The playground quotes the exact cost before anything runs.

It is the AI kind. A plain converter puts a still image or a waveform behind your MP3 so you can upload it to YouTube, and free editors like CapCut or VEED do that well. Percify makes a face speak your audio, which is what you want when the video should look like someone talking.

Yes. Upload the MP3 with a photo and you get an MP4 of the photo speaking it. Other common audio files such as WAV and M4A work too.

The fast lip sync model costs 2 credits a second of video, so a 15 second clip is 30 credits and a 60 second clip is 120. The higher quality model is 4 credits a second at 480p or 6 at 720p. Runs that fail are refunded.

On the fast engine the median 17 second clip takes about 86 seconds to render. The higher quality engine takes longer.

Yes. Any speech works best when one person is talking at a time and the recording is clean. Cut a long episode into the clip you want to share, because you pay per second of video.

Yes. Pick a voice and type the script, or clone your own voice from a short clean sample at 2 credits a minute of audio, then make the photo speak it. Only clone a voice you have the rights to.

Your own, or someone who has agreed to it. Percify’s terms require that you hold the rights to what you upload. Do not upload a public figure or anyone who has not agreed.

No plan adds a watermark, and under the terms you own the content you create, so you can use it in ads, courses and client work. Output is not exclusive.