Quick Answer
A clone of you in 2026 is a recording engine, not a copy of you. From one photo and 15 to 30 seconds of speech, samples from about 5.6 seconds being accepted, it reproduces your timbre and much of your cadence, and it will read any script you give it at 2 credits a second, about $2.53 a minute on Percify's Creator plan. What it does not do is think, improvise, hold a conversation or know anything you have not written down. It also sounds least like you exactly where you sound most like yourself: laughing, interrupting, losing the thread and finding it again.
What a face and voice clone reproduces from 15 to 30 seconds of speech, what it cannot do (think, respond, improvise), where it convinces, and what it costs. Checked 15 September 2026.
Related next steps
The interesting question is not whether an AI version of you is possible, it is how close it gets and where it stops. This page answers both, with the specifications and the costs, read on 15 September 2026.
What it needs, stated correctly
One photo where your face fills at least a third of the frame, and 15 to 30 seconds of clear speech. Samples from about 5.6 seconds are accepted, though a longer one carries more of your voice.
That is worth saying plainly because this site previously told readers that cloning needs "typically a few minutes of audio". It does not, and a page that overstates the input is telling you to spend time you do not need to spend.
What it reproduces
- Timbre, the quality that makes a voice recognisably yours on a phone call.
- Much of your cadence, the pace and the pauses in the sample you gave it.
- Your face in motion, synced to whatever audio you supply.
For a scripted 30 to 60 seconds, that is usually enough for someone who knows you to accept it without comment. Clone Yourself makes both for 5 credits each, once, which the free plan's 10 credits cover, and InfiniteTalk Fast ↗ renders at 2 credits a second, so a minute is 120 credits, about $2.53 on Creator.
What it does not do, which is the useful part
- It does not think. There is no conversational agent behind it and nothing that answers a question. It reads what you wrote.
- It does not respond, live or otherwise. Percify has no live avatar you can talk to.
- It does not improvise, so the delivery is as good as the script and no better.
- It does not know anything about you, including the things you would normally mention.
And the honest limit on resemblance: a clone sounds least like you where you sound most like yourself. Laughter, an interrupted sentence, a thought abandoned halfway and picked up again, the moment you get genuinely annoyed. Scripted competence is what these models do well.
Where that leaves it useful
Short scripted segments where the words matter more than the performance: an intro, an update, an explainer, a version of the same message in another language, or a video you would otherwise not have made at all because setting up a camera was the obstacle.
Where it does not belong: anything emotional, anything reactive, and anything where being visibly present is the point. If you would notice the difference, your audience will.
Consent and disclosure, briefly
Your own face and voice, and you decide. Someone else's needs written permission naming AI cloning, because Tennessee's ELVIS Act treats a simulation of a voice as the voice. Realistic AI video must be disclosed: the EU AI Act has required it since it began applying on 2 August 2026, and TikTok, YouTube and Meta each have their own label. Details in voice cloning and the law and the four duties.
For the tools, including ones you can run yourself, see cloning your voice free; other avatar platforms are in the best AI avatar generators compared; prices are on pricing.
Ready to Create Your Own AI Avatar?
Make a talking video of yourself from one photo, in your own voice. No watermark on any plan.
Get startedCancel anytime, no contracts
Got questions?
Frequently asked
On Percify, 15 to 30 seconds of clear speech, with samples from about 5.6 seconds accepted. Claims that it needs "a few minutes of audio" overstate it, including one that previously appeared on this site.
Not on Percify. It renders recorded video from a script or an audio file. There is no conversational agent, no live avatar and nothing that responds to questions. A clone is a recording engine, not an assistant.
For scripted speech of 30 to 60 seconds, usually convincing to people who know you. It reproduces timbre and much of your cadence. It is least convincing exactly where you are most yourself: laughing, interrupting, losing your thread, or genuinely reacting.
Read on 15 September 2026: 5 credits for the face and 5 for the voice, once, which the free plan's 10 credits cover, then 2 credits a second of video, so a minute is 120 credits, about $2.53 on the Creator plan ($26 for 1,233 credits).
For realistic content, yes. The EU AI Act has required deepfake content to be disclosed since it began applying on 2 August 2026, and TikTok, YouTube and Meta each require their own label or setting. Read on 16 September 2026.
