Quick Answer
how toA UGC-style video ad on Percify is assembled from four steps: find a format that already works, script for roughly twenty seconds, render it as a talking avatar in 4:5, and cut b-roll around it. The length is not a style choice — across 179 finished renders the median clip people keep is 19 seconds and 66% are under 30. Competitor ad creative can be pulled in through a market scan's Ad Library leg, but that source returns no view counts, so it shows what rivals pay to run rather than what works.
A UGC-style ad is four decisions, not one button. The format, the length, the voice and the proof — and what competitor ad data can and cannot tell you.
Keep reading
Related next steps

There is no button that turns a product into a converting advert, here or anywhere else. What there is: a short set of decisions, each of which has a right answer more often than people assume.
- What format? Which structure you are performing.
- How long? Shorter than you think.
- Whose voice and face? Yours, or a presenter's.
- What proof? The thing that makes the claim believable.
Everything below is about answering those four with real numbers rather than taste.

Length: the number that surprises people
We looked at every finished talking-head render on the platform between 7 May and 2 September 2026 — 179 renders across 61 accounts:
| measure | value |
|---|---|
| median clip | 19.0 seconds |
| under 30 seconds | 66% |
| over two minutes | 2.8% |
That is what people actually finish and publish, against an industry norm of recommending a two-minute case study. Under 3% of real output is that long.
For an ad specifically, write for twenty seconds: one line of hook, two or three of substance, one of exit. Read it aloud before rendering — if you run out of breath, the model will faithfully deliver a rushed performance rather than fixing it for you.
Format: 4:5 unless you know otherwise
Vertical-ish 4:5 is the workhorse ad ratio. It occupies more of a feed than 1:1, survives being cropped to 1:1, and does not get letterboxed the way full 9:16 does when a platform decides to show it in-feed. Google's video ad specifications ↗ draw the same distinction between feed and full-screen placements. Meta's own placement guidance ↗ and TikTok's creative specs ↗ both point the same way for feed placements.
Getting this wrong is the most common avoidable mistake in AI-made ads, and it is entirely mechanical: the complete guide to 4:5 video dimensions covers the safe areas and the export settings, and the common mistakes version covers what breaks.
The pieces that build it
On Percify, a UGC-style ad is assembled from studios that each do one thing well:
| step | what to use |
|---|---|
| the spokesperson clip | Avatar Studio — one photo, one script |
| the voice | Voice Studio — cloned, or one of 52 presets |
| a proven structure | Content Replication — paste a short that worked |
| b-roll and product plates | Realtime Video — 5 to 15 second clips, fast |
| the same ad in another language | Voice Studio's dub, with the mouth regenerated from the translated track |
The order that works: find the format first, write to it, then render. Writing a script and hunting for a structure afterwards is how you end up with a monologue in a feed full of edits.

What competitor ad data actually tells you
A brand market scan on Percify has an opt-in leg that pulls competitor creative from the Meta Ad Library ↗. It costs extra because it is genuinely the expensive half of a scan.
It is worth being precise about what it is good for, because this is where ad research usually goes wrong:
So use it to see creative and claims, not to rank them. For ranking, the organic side of a scan carries real view counts, and that is covered in Viral Intelligence.
Proof is the part AI cannot invent
The hook and the cutting can be replicated. The proof cannot, and it is the half that decides whether the ad works.
Proof is a number you actually have, a before-and-after you actually ran, a screenshot of a real message, or a specific outcome for a specific customer. Everything else is a claim, and audiences have been trained by a decade of performance marketing to discount claims.
This matters more, not less, when the presenter is synthetic. A generated face making an unbacked assertion is exactly the thing viewers have learned to scroll past. A generated face reading a real number is a different object. What to check before you publish covers the disclosure side of the same question.

A workable sequence
- Find three ads in your category that have been running a while. Long-lived beats clever.
- Pick one structure and paste it into Content Replication, or write to its shape by hand.
- Write twenty seconds. Hook, substance, exit. Put the proof in the substance.
- Render at 4:5, 480p, with your own cloned face and voice if the ad is about you or your business.
- Add b-roll with fast 5-second generations rather than stock footage — a plate that matches your product beats a library clip that does not.
- Make the second version before you judge the first. Single-ad tests tell you nothing.
For the same ad in another market, dub it — but regenerate the mouth from the translated audio, or the shapes belong to the wrong language and it reads as fake for reasons nobody can name. That failure is described in full in why AI lip sync looks wrong.
The full feature list shows every studio in one place, pricing is one credit meter across images, voices and video, and the explore page is real output rather than a showreel.
Related on Percify
- Get 4:5 right — the safe areas and export settings.
- Common 4:5 mistakes — what breaks.
- Testimonial videos with AI — the proof-shaped format.
- Viral Intelligence — see what rivals are running.
- Every studio at a glance — the whole product on one page.
- What it costs — one credit meter.
- Real output — finished work rather than a showreel.
Ready to Create Your Own AI Avatar?
Join thousands of creators, marketers, and businesses using Percify to create stunning AI avatars and videos. Start your free trial today!
Get Started FreeGot questions?
Frequently asked
Around twenty seconds. Across 179 finished renders on Percify the median clip people keep is 19.0 seconds, 66% are under 30 seconds and only 2.8% run past two minutes. Structure it as one line of hook, two or three of substance and one of exit.
4:5 for feed placements. It occupies more of the feed than 1:1, survives a crop to 1:1, and avoids the letterboxing full 9:16 can get when a platform shows it in-feed.
Yes. A spokesperson clip needs one photo and a script, the voice can be cloned from a short sample or taken from 52 presets, and b-roll can be generated in 5 to 15 second clips. What AI cannot supply is the proof — the real number or outcome that makes the claim believable.
Not directly. The Facebook Ad Library returns no view counts, so its creative cannot be ranked by performance. It shows what a competitor is paying to run and how long it has run, which is a weaker signal than views. Organic competitor data carries real view counts and is the better ranking source.
Dub it in Voice Studio, and regenerate the mouth movement from the translated audio rather than the original. Visible mouth shapes differ between languages, so reusing the original animation makes a correctly translated ad read as fake.
