Turn text into a voice

Paste a script, pick a voice or clone your own, and download the audio. Twelve voice models in one place, from 2 credits a clip, no watermark on any plan, and you own the result.

Generate a voice

Every voice model, and what it costs

You pay per clip you generate rather than per seat or per month. Most of these charge a flat price per run, so a long script costs the same as a short one; the two that scale with length say so. Runs that fail are refunded.

ModelPriceAboutGood for
Qwen3 TTS Flash2 credits a clip$0.03 to $0.05The fastest and cheapest, English and Chinese
Zonos 22 credits a minute of audio$0.03 to $0.05Clone a voice from a short sample, multilingual
Seed Speech TTS 2.03 credits a clip$0.05 to $0.08Natural delivery across many languages
Gemini TTS4 credits per 1,000 characters$0.07 to $0.10Expressive, and two speakers in one take
Speech-02-HD5 credits a clip$0.08 to $0.13Emotional range when the line has to land
Speech-02-Turbo5 credits a clip$0.08 to $0.13The same voice engine, faster
Chatterbox Turbo5 credits a clip$0.08 to $0.13The quickest open source model we run
ElevenLabs Multilingual V210 credits a clip$0.17 to $0.25The voice most people mean by "AI voice"
ElevenLabs Eleven V310 credits a clip$0.17 to $0.25ElevenLabs newest, for the take you keep
XTTS-v210 credits a clip$0.17 to $0.25Multilingual cloning, open weights
Speech 2.8 HD10 credits a clip$0.17 to $0.25Clean pronunciation on long scripts
Seed Audio 1.030 credits a clip$0.50 to $0.75Speech plus the sound around it, from one prompt

Credits work out at about 2 cents each depending on the plan: 20 credits is $0.33 on Starter, $0.42 on Creator and $0.50 on a credit pack. Plan prices are on the pricing page. Checked 22 September 2026.

How it works

  1. Step 1

    Paste the script

    Type or paste what the voice should say. Punctuation does real work here: it is what the model reads as pauses and breath.

  2. Step 2

    Pick or clone a voice

    Use a ready made voice, or clone one from a short clean sample of someone who agreed to it. Check the sample below first.

  3. Step 3

    Generate and download

    Get an MP3 or WAV back, with no watermark. Try the same script on two models before you settle on one.

Check your voice sample first, free

A clone is only as good as the recording it came from. This rates length, background noise, clipping and silence, and it runs entirely in your browser. Nothing is uploaded.

Is this recording good enough to clone?

Record 20 to 30 seconds or choose a file. The checker measures length, distortion, level, background noise and silences in your browser; nothing is uploaded.

Which voice model to use

  • Testing a script. Qwen3 TTS Flash at 2 credits. You can hear the same line eight ways for the price of one premium take.
  • The voice people picture. ElevenLabs Multilingual V2 and Eleven V3, 10 credits a clip.
  • A line that has to carry feeling. Speech-02-HD holds emotion where flatter models level it out.
  • Your own voice. Zonos 2 at 2 credits a minute, or XTTS-v2. Both clone from a short sample and both speak several languages.
  • Two people in one take. Gemini TTS handles two speakers in a single generation, charged per 1,000 characters.

Cloning a voice, and whose voice you may clone

Cloning works from a short, clean sample: one person, no music, no room echo. Save it once and every later script is read in that voice. Cloning through these models is multilingual, so a voice recorded in one language can speak another.

  • Use your own voice, or one whose owner has agreed to it. A person's voice is theirs.
  • Do not upload a recording of a public figure. Percify's terms require you to hold the rights to what you upload.
  • You own what you generate, under those terms. It is not exclusive.
  • No watermark on the audio, on any plan including the free one.

Put the voice on a face

Most people generating a voice want it coming out of something. Once you have the audio, lip sync it to a photo from 2 credits a second and the mouth, jaw and head move with it. If you want a picture to move without speech, that is image to video.

Hear your script out loud

Twelve voice models, from 2 credits a clip, no watermark, and failed runs refunded.

Generate a voice

Frequently asked questions

Everything you need to know before you start.

It is a model that reads text aloud in a synthetic voice. You type or paste a script, pick a voice, and get an audio file back. Some models also clone a voice: give them a short sample of someone speaking and they read your script in that voice.

From 2 credits a clip on Percify, which is about 3 to 5 cents. ElevenLabs Multilingual V2 and Eleven V3 are 10 credits, roughly 17 to 25 cents. Gemini TTS is charged by length instead, at 4 credits per 1,000 characters. Credits work out at about 2 cents each depending on the plan, and runs that fail are refunded.

It depends on the job: ElevenLabs Multilingual V2 is what most people mean by a realistic AI voice, Speech-02-HD carries emotion better, Qwen3 TTS Flash is the cheapest and fastest at 2 credits, and Zonos 2 or XTTS-v2 are the ones that clone a voice. Percify runs all twelve on one balance, so you can try the same script on several. It is our product, so compare it against the tools you are considering on the same points.

Percify has a free plan with no watermark, and the voice sample checker on this page is free and runs entirely in your browser. Generating speech itself costs credits on every plan, because every run costs us provider time. The honest version is that you can start without paying and hear real output before deciding.

Yes. Record or upload a clean sample and Zonos 2 or XTTS-v2 will read new scripts in that voice, in several languages. The free checker on this page rates the sample for length, noise, clipping and silence first, because a clone is only ever as good as the recording it came from.

Only with their permission. Percify’s terms require that you hold the rights to whatever you upload, and a person’s voice is theirs. Do not upload a recording of a public figure or of anyone who has not agreed to it.

Yes. Zonos 2 and XTTS-v2 are both multilingual voice cloning models, so a cloned voice can speak a language the original sample was not in. One thing to know: the one-click clone inside Voice Studio generates English, while the playground models are the multilingual route.

Yes. Percify’s terms say you own the AI generated content you create, subject to the terms. Output is not exclusive, and you need the rights to any voice you clone.

Yes, and that is the usual reason people generate one. Once you have the audio, lip sync it to a photo so the face speaks it, which runs from 2 credits a second. If you want the picture to move without speech instead, that is image to video.

MP3 or WAV, downloadable, with no watermark on any plan including the free one.

It depends on the model. The flat priced models charge per run rather than per second, so a longer script costs the same as a short one on those. Zonos 2 is charged per minute of audio and Gemini TTS per 1,000 characters, so those two scale with length. The playground quotes the exact cost before you run anything.

Yes. All twelve models run through the same API with one key and one credit balance, so changing voice model is a parameter rather than a new integration. Details are on the API pricing page and in the docs.