Podcast Studio · two speakers

An AI podcast where both speakers listen

Percify's AI podcast generator turns a written conversation into one video of two people talking, where each speaker animates from their own audio track and visibly listens while the other speaks. You write the dialogue and pick the voices, presets or your own clone; there is nothing to record. It costs 150 credits per minute, runs from 15 seconds to a 3 minute cap, and renders at 480p.

Cancel anytime, no contracts
150
credits per minute of video
15 s
shortest podcast
3:00
longest podcast
2
speakers, each listening on camera

Credits cost about 2 cents each depending on the plan, so a minute of talking video at 120 credits works out at roughly $2 to $3.

White line illustration for Percify Podcast Studio, a two person AI podcast

Why most AI podcasts look like two monologues

The model underneath takes two audio tracks and an order. Read literally, that means speaker A for ninety seconds, then speaker B for ninety seconds: one person talks while the other sits frozen, then they swap. Most tools that promise two speakers ship exactly that, and it shows.

How the conversation is actually built

Every turn is synthesized separately. Then both audio tracks are built at full length, with silence wherever the other person is talking, and the two are rendered together rather than one after the other.

Each side animates from the energy of its own track, so a stretch of silence is not a gap: it renders as a person listening. That was checked on a real render before the feature shipped, because the failure is subtle and looks fine in a still frame.

Line drawing of two interleaved waveforms, where one is loud the other is flat, two speakers taking turns

You write it, you do not record it

You write the dialogue and the voices are synthesized. That changes the work: you can draft a conversation, read it, and change a line, which is a different activity from editing a recording, and scripts get better when revising them is cheap.

Put your own voice and face on one side with clone yourself. If you already have recorded audio, the API accepts two tracks directly.

What it costs, and why there is no 720p

The arithmetic is published because it explains the price. At 720p the provider charges $0.06 a second, and the same 150 credits would return a 4% margin, so HD would mean roughly doubling the price. That is a pricing decision, not a missing setting.

InputValue
Provider rate at 480p$0.03 per second
One minute of output, to us$1.80
Our minimum margin40%
Minimum defensible price120 credits
What we charge150 credits a minute

Why three minutes is the ceiling

The engine spends 10 to 30 seconds of compute per second of output, so three minutes of video is 30 to 90 minutes of waiting, and ten minutes could take five hours. For a long form show, make several segments. For a two minute conversation that explains something, this is built for exactly that.

One more detail, stated rather than rounded away: the provider always renders exactly one second more than the audio it is given, and that second is included in the billed length because we are charged for it.

What to make with two speakers

Two voices suit different material from a single talking head:

  • An objection and the answer to it.
  • An interview, with you as the host and a preset voice as the guest.
  • A short explainer where one voice asks the questions a listener would.
  • A before and after conversation about a product or a decision.
Line drawing of two figures facing each other joined by one line, one speaking while the other listens

Make an episode in four steps

  1. 1

    Write the conversation

    Short turns read more naturally than long speeches.

  2. 2

    Pick two voices

    Two distinct presets, or your own clone on one side.

  3. 3

    Choose the length

    Anything from 15 seconds to 3 minutes, at 150 credits a minute.

  4. 4

    Render and publish

    One video, both speakers on screen, each listening while the other talks.

Podcasts with no camera

Real Percify output: podcast style hosts made from one photo, voices from a script.

Frequently asked questions

Can AI generate a video podcast from a script?

Yes. Write the dialogue and Percify renders one video of two people talking, where each speaker animates from their own track and listens while the other speaks.

How much does an AI podcast cost?

150 credits per minute on Percify, from 15 seconds up to a 3 minute cap, rendered at 480p.

Can one of the speakers be me?

Yes. Clone your face from one photo and your voice from a sample of at least 5.6 seconds, 5 credits each, once. Then take one side of the conversation yourself.

Why is the maximum three minutes?

The engine needs 10 to 30 seconds of compute per second of output, so three minutes already means 30 to 90 minutes of waiting. Longer shows work better as several segments.

Can I upload recorded audio instead of writing a script?

In the app you write the dialogue and the voices are synthesized. The public API accepts two audio tracks for callers who already have recordings.

Keep reading

Write your first episode

Two voices, one script, fifteen seconds to three minutes.

Cancel anytime, no contracts