Podcast Studio · two speakers
An AI podcast where both speakers listen
Percify's AI podcast generator turns a written conversation into one video of two people talking, where each speaker animates from their own audio track and visibly listens while the other speaks. You write the dialogue and pick the voices, presets or your own clone; there is nothing to record. It costs 150 credits per minute, runs from 15 seconds to a 3 minute cap, and renders at 480p.
- 150
- credits per minute of video
- 15 s
- shortest podcast
- 3:00
- longest podcast
- 2
- speakers, each listening on camera
Credits cost about 2 cents each depending on the plan, so a minute of talking video at 120 credits works out at roughly $2 to $3.

Why most AI podcasts look like two monologues
The model underneath takes two audio tracks and an order. Read literally, that means speaker A for ninety seconds, then speaker B for ninety seconds: one person talks while the other sits frozen, then they swap. Most tools that promise two speakers ship exactly that, and it shows.
How the conversation is actually built
Every turn is synthesized separately. Then both audio tracks are built at full length, with silence wherever the other person is talking, and the two are rendered together rather than one after the other.
Each side animates from the energy of its own track, so a stretch of silence is not a gap: it renders as a person listening. That was checked on a real render before the feature shipped, because the failure is subtle and looks fine in a still frame.

You write it, you do not record it
You write the dialogue and the voices are synthesized. That changes the work: you can draft a conversation, read it, and change a line, which is a different activity from editing a recording, and scripts get better when revising them is cheap.
Put your own voice and face on one side with clone yourself. If you already have recorded audio, the API accepts two tracks directly.
What it costs, and why there is no 720p
The arithmetic is published because it explains the price. At 720p the provider charges $0.06 a second, and the same 150 credits would return a 4% margin, so HD would mean roughly doubling the price. That is a pricing decision, not a missing setting.
| Input | Value |
|---|---|
| Provider rate at 480p | $0.03 per second |
| One minute of output, to us | $1.80 |
| Our minimum margin | 40% |
| Minimum defensible price | 120 credits |
| What we charge | 150 credits a minute |
Why three minutes is the ceiling
The engine spends 10 to 30 seconds of compute per second of output, so three minutes of video is 30 to 90 minutes of waiting, and ten minutes could take five hours. For a long form show, make several segments. For a two minute conversation that explains something, this is built for exactly that.
One more detail, stated rather than rounded away: the provider always renders exactly one second more than the audio it is given, and that second is included in the billed length because we are charged for it.
What to make with two speakers
Two voices suit different material from a single talking head:
- An objection and the answer to it.
- An interview, with you as the host and a preset voice as the guest.
- A short explainer where one voice asks the questions a listener would.
- A before and after conversation about a product or a decision.

Make an episode in four steps
- 1
Write the conversation
Short turns read more naturally than long speeches.
- 2
Pick two voices
Two distinct presets, or your own clone on one side.
- 3
Choose the length
Anything from 15 seconds to 3 minutes, at 150 credits a minute.
- 4
Render and publish
One video, both speakers on screen, each listening while the other talks.
Podcasts with no camera
Real Percify output: podcast style hosts made from one photo, voices from a script.
Mollie
a podcast host
Podcast
a podcast clip, no camera
Alex
talks from one photo
Frequently asked questions
Can AI generate a video podcast from a script?
Yes. Write the dialogue and Percify renders one video of two people talking, where each speaker animates from their own track and listens while the other speaks.
How much does an AI podcast cost?
150 credits per minute on Percify, from 15 seconds up to a 3 minute cap, rendered at 480p.
Can one of the speakers be me?
Yes. Clone your face from one photo and your voice from a sample of at least 5.6 seconds, 5 credits each, once. Then take one side of the conversation yourself.
Why is the maximum three minutes?
The engine needs 10 to 30 seconds of compute per second of output, so three minutes already means 30 to 90 minutes of waiting. Longer shows work better as several segments.
Can I upload recorded audio instead of writing a script?
In the app you write the dialogue and the voices are synthesized. The public API accepts two audio tracks for callers who already have recordings.
Keep reading
Write your first episode
Two voices, one script, fifteen seconds to three minutes.