Which voice model to use
- Testing a script. Qwen3 TTS Flash at 2 credits. You can hear the same line eight ways for the price of one premium take.
- The voice people picture. ElevenLabs Multilingual V2 and Eleven V3, 10 credits a clip.
- A line that has to carry feeling. Speech-02-HD holds emotion where flatter models level it out.
- Your own voice. Zonos 2 at 2 credits a minute, or XTTS-v2. Both clone from a short sample and both speak several languages.
- Two people in one take. Gemini TTS handles two speakers in a single generation, charged per 1,000 characters.
Cloning a voice, and whose voice you may clone
Cloning works from a short, clean sample: one person, no music, no room echo. Save it once and every later script is read in that voice. Cloning through these models is multilingual, so a voice recorded in one language can speak another.
- Use your own voice, or one whose owner has agreed to it. A person's voice is theirs.
- Do not upload a recording of a public figure. Percify's terms require you to hold the rights to what you upload.
- You own what you generate, under those terms. It is not exclusive.
- No watermark on the audio, on any plan including the free one.
Put the voice on a face
Most people generating a voice want it coming out of something. Once you have the audio, lip sync it to a photo from 2 credits a second and the mouth, jaw and head move with it. If you want a picture to move without speech, that is image to video.