Percify Clone Yourself: Build a Digital Twin

Percify Team

Percify Team

Content Writer

September 5, 2026
7 min read
Clone Yourself Ai / Build A Digital Twin Avatar / Ai Version Of Me For Video

Quick Answer

how to

Clone Yourself builds the two assets every other Percify studio consumes: a face, from one clear photo, and a voice, from a sample of at least 5.6 seconds. A custom avatar costs 5 credits and a voice clone costs 5 credits, both charged once. It is a setup step rather than a studio — once both exist, the talking avatar, the podcast, the dub and the replicated short all run as you instead of as a stock presenter.

Clone Yourself is a setup step, not a studio. What a good source photo and a good voice sample look like, what each costs, and why doing it first changes everything.

The Percify Clone Yourself title lockup, white and gold three dimensional lettering glowing against a black background

Every other studio on Percify asks the same two questions before it can do anything: whose face, and whose voice. Clone Yourself answers both, once, and then gets out of the way.

That framing matters because it changes the order you should work in. People usually arrive wanting a video, spend their first credits on a stock presenter, decide the result is fine but impersonal, and only then wonder whether they could have used themselves. Doing the clone first costs about ten credits and makes everything afterwards yours.

A white line drawing on black of a single photograph and a short waveform converging into one figure, representing a face and a voice becoming one avatar

The face: one photo, and what makes it good

You upload one clear photo. There is no scanning session, no turntable and no video capture.

Because a lip-sync model rebuilds the mouth region rather than animating a rig, the photo decides quality more than any setting does. The rules are unglamorous:

- The face should fill at least a third of the frame. Below roughly a quarter of frame width there is not enough mouth detail to rebuild, and the result reads soft no matter what resolution you render at.

- Nothing across the mouth. No hand, no microphone, no stray hair.

- Even, frontal light. Hard side light bakes a shadow into every frame the model generates from it.

- A neutral or slightly open mouth. A wide smile is a hard starting position to animate out of.

About 9% of clips get retried on the platform, and the overwhelming majority of those retries trace back to the source image rather than to the render. Choosing the photo carefully is the cheapest quality decision available — the full diagnosis of what goes wrong is worth five minutes before you upload.

A custom avatar costs 5 credits. Generating from it afterwards costs 2.

The voice: 5.6 seconds is the floor, not the target

The voice clone takes a reference recording. The engine needs at least 5.6 seconds of it; anything shorter is padded by repeating the clip to about 6.2 seconds so the request still works.

That padding is a courtesy, not a feature. A three-second clip looped twice still contains three seconds of vocal information, and the clone comes out flat. Record 15 to 30 seconds and say something with range in it — a question, a statement, a number — rather than one even sentence.

Room matters more than microphone. Curtains, a sofa or a bed beat bare walls and a good mic. Keep a steady distance. If your levels swing, normalise the file first with ffmpeg's loudnorm filter ↗.

Cloning costs 5 credits, once. The voice side is covered in full here.

A white line drawing on black of a microphone with a measuring scale beside it marking a minimum threshold, representing a sample length floor

What the twin unlocks

Once a face and a voice exist, they are inputs everywhere:

studiowhat your twin does there
Avatar Studiosays a written script to camera
Voice Studiospeaks a dub in your own voice
Podcast Studiotakes one side of a two-person conversation
Content Replicationperforms a proven short's structure as you

That table is the argument for doing this first. Platforms that build the presenter for you, such as Synthesia ↗, keep the likeness inside their own product; a twin you own is one you can point at any of these.

That last one is the reason to bother. A format that worked for someone else, performed by a stock presenter, is a copy. Performed by you, it is your version of a structure that already earns attention. The replication write-up explains what actually gets reused and what does not.

The part worth thinking about before you start

A digital twin is a likeness, and likenesses carry obligations. Two are practical rather than philosophical:

  1. Clone only faces and voices you have the right to use. Yours is the easy case. A colleague, a client or a public figure is not, regardless of what the tool will technically accept. The EU AI Act's transparency provisions ↗ and equivalent rules elsewhere are converging on disclosure for synthetic likenesses, and the direction of travel is one way. The industry answer taking shape is provenance metadata: the C2PA standard ↗ attaches a signed record of how a file was made, and the major camera and model vendors are adopting it.
  2. Decide your disclosure position once. Whether you label AI-presented video is a brand decision, not a legal minimum in every market — but deciding it deliberately beats deciding it after someone asks. We wrote up what to check before publishing as a checklist rather than an essay.

None of that is a reason not to build a twin. It is a reason to build one of yourself.

What it costs, end to end

stepcreditshow often
custom avatar5once
voice clone5once
avatar generation2per generation
talking-head render (480p average)69per clip

The setup is roughly a tenth of what one finished video costs, which is why doing it first is close to free in practice. A free account starts with credits and no card, so the twin can exist before you decide anything about plans — the full pricing is one meter across images, voices and video, with no quality or watermark tiers to buy past.

A white line drawing on black of two small tokens on the left of a scale and one large token on the right, showing a small setup cost against a larger per-video cost

Do this in fifteen minutes

  1. Find a photo where your face fills a third of the frame, evenly lit, mouth clear.
  2. Record 20 seconds of yourself talking in a soft room at a steady distance.
  3. Build the avatar (5 credits) and the voice (5 credits).
  4. Render one 20-second clip at 480p to check the pair works together.
  5. Only then decide what you actually want to make.

Step 4 is the one people skip. A short test render tells you within three minutes whether the photo was the right photo, and it costs less than discovering it on a script you cared about. If you would rather see other people's results first, the explore page is real output rather than a showreel, and the comparison against other avatar generators is a fair place to start if you are not sure Percify is the tool.

- Avatar Studio — make the twin speak.

- Voice Studio — the voice half in detail.

- Content Replication — perform a proven format as yourself.

- Podcast Studio — put the twin in a conversation.

- Every studio at a glance — the whole product on one page.

- Lip sync tools compared — what else is out there.

- Real output — results rather than a showreel.

Ready to Create Your Own AI Avatar?

Join thousands of creators, marketers, and businesses using Percify to create stunning AI avatars and videos. Start your free trial today!

Get Started Free

Got questions?

Frequently asked

One clear photo where your face fills at least a third of the frame, and a voice recording of at least 5.6 seconds — ideally 15 to 30 seconds. No filming session, no scanning rig and no video capture.

About 10 credits: 5 for the custom avatar and 5 for the voice clone, both charged once. Generating from the avatar afterwards costs 2 credits, and a finished talking-head render averages 69 credits at 480p.

Yes, provided the framing is right. Lip-sync models rebuild the mouth region from the source frame, so a close-up with even light and an unobstructed mouth matters far more than the number of photos.

Before. The setup is roughly a tenth of what a single finished video costs, and every studio consumes the face and the voice. Making the first video with a stock presenter usually means making it twice.

Cloning your own likeness is straightforward. Cloning someone else's — a colleague, a client, a public figure — needs their permission, and transparency rules such as the EU AI Act's provisions on synthetic media are moving toward required disclosure. Decide your labelling position deliberately rather than after someone asks.

clone yourselfdigital twinai avatarvoice cloningpercify
Percify Team
Published on
Share article

Related Reads

Stop Paying for HeyGen: Percify's 2026 AI Avatar Breakthrough for Content Scaling - Percify AI Avatar Blog Cover
Ai Avatar For Content ScalingJul 24, 26

Stop Paying for HeyGen: Percify's 2026 AI Avatar Breakthrough for Content Scaling

Unlock massive content scaling with Percify's AI avatar technology. Get photorealistic videos in 140+ languages for <$0.25/min. Compare 2026 pricing vs. HeyGen ($48/mo) & Synthesia.

Read Article
5 AI Avatar Videos for Agencies: Stop Overpaying for Voice Clone Until You Read This (July 2026) - Percify AI Avatar Blog Cover
Ai Avatar For Content AgenciesJul 20, 26

5 AI Avatar Videos for Agencies: Stop Overpaying for Voice Clone Until You Read This (July 2026)

Struggling with expensive AI avatar videos? Discover 5 ways agencies can leverage voice cloning with Percify for under $0.25/min. Compare pricing & features.

Read Article
Unmasking the 3 Biggest AI Avatar Generator Mistakes You're Still Making in 2026 - Percify AI Avatar Blog Cover
Avatar GeneratorJun 30, 26

Unmasking the 3 Biggest AI Avatar Generator Mistakes You're Still Making in 2026

Frustrated with robotic AI avatars or slow video generation? Discover how to create photorealistic AI avatar videos with perfect lip-sync and voice cloning in 140+ languages, fast. Compare top tools.

Read Article
The 3 Hidden Costs of AI Video You Missed in 2026 (Percify vs Colossyan) - Percify AI Avatar Blog Cover
Percify Vs Colossyan For EnterpriseJun 24, 26

The 3 Hidden Costs of AI Video You Missed in 2026 (Percify vs Colossyan)

Unlock advanced AI video for $6.99/mo with Percify, 7x less than HeyGen. Discover 3 critical enterprise differences in percify vs colossyan for enterprise in 2026 and save thousands.

Read Article
Synthesia Alternative: Can I Use a Photo as an Avatar in 2026? Percify vs. Competitors - Percify AI Avatar Blog Cover
Can I Use A Photo As An AvatarJun 21, 26

Synthesia Alternative: Can I Use a Photo as an Avatar in 2026? Percify vs. Competitors

Wondering 'can I use a photo as an avatar'? Discover how Percify creates photorealistic AI videos from your photos and voice, costing just ~$0.25/min. Compare with Synthesia, HeyGen & more.

Read Article
Percify Voice Studio: Clone a Voice, Then Dub - Percify AI Avatar Blog Cover
Percify Voice Studio / Clone Your Voice For Video / Ai Dubbing VoiceSep 5, 26

Percify Voice Studio: Clone a Voice, Then Dub

How voice cloning and dubbing actually work on Percify: the 5.6 second sample floor, 52 preset voices, what cloning costs, and the trap that breaks dubbed lip sync.

Read Article

Create anywhere with Percify

Try Percify for free, and explore all the tools you need to create, voice, and animate your digital avatars.

Start free then upgrade as you grow.