Quick Answer
how toTo remake a video shot for shot, the source is analysed to find every cut, split at those boundaries, and each shot regenerated before reassembly. Across 135 real replications a 34.2 second source became about 14 clips averaging 2.9 seconds, for roughly 4.8 credits. One flow in five fails, and nearly nine in ten of those failures happen at the analysis step rather than in generation, which is why uploading the file instead of pasting a link matters: uploads failed 0 percent against up to 28.6 percent for links.
Measured across 135 replications and 843 clips: shortform cuts every 2.9 seconds, one flow in five fails, and uploads fail 0 percent while links fail up to 29.
Keep reading
Related next steps

You find a video that works. The format is obvious, the pacing is obvious, and you still have no way in, because what you are looking at is forty edit decisions somebody already made and did not write down.
This is about closing that gap: taking a video you like and producing your own version of it, shot for shot, with the cuts in the same places. It is the most used thing we run, and the numbers below come from 135 real replications and 843 generated clips, not from a feature page. Where something fails, the failure rate is stated. Where a claim is about another platform's rules, it links to that platform.
What actually happens to the video
The pipeline is three steps, and only the middle one is interesting. Each regenerated shot is produced by a specific engine, and the model catalogue records which one, so a result can be traced rather than guessed at.
First the source is analysed: every cut is found, and the video is split at those boundaries. Then each shot is described and regenerated as your version. Then the pieces go back together in order.
The first step is the one that decides whether the whole thing works, and it is also where almost every failure lives. More on that below, because it is the practical part.
What a real replication looks like
Averaged across those 135 flows: The same records feed the flow recipes, which is how a format gets pinned to a pipeline once it works.
| measure | value |
|---|---|
| source video length | 34.2 seconds |
| clips it becomes | 13.8 |
| average clip length | 2.9 seconds |
| credits spent | 4.8 |
| flows that failed | 20.7 percent |
The number worth sitting with is 2.9 seconds. A shortform video that holds attention cuts roughly every three seconds, and a 34 second video is carrying about fourteen separate shots.
Most people trying to make this kind of video by hand cut far less often, and that single difference is usually why their version feels slow next to the original. It is not the colour, it is not the music, it is the cutting rhythm. You can see the cost of that many shots in the credit pricing before committing.
Cutting rhythm is the format
If you take one thing from this page: the format you are copying is mostly a pacing decision, and pacing is measurable. Count the cuts in the video you like, divide by its length, and you have the spec. Editors have measured this for decades: the average shot length ↗ of mainstream film has been falling since the 1950s, and shortform simply arrived at the end of that curve rather than inventing it.

Under three seconds a shot, the video reads as fast and modern. Above about five, it reads as a talking video that happens to have edits. Neither is wrong, but they are different formats and mixing them is what makes a video feel unresolved.
This is why replicating shot boundaries works better than replicating a description. A prompt that says "fast paced and engaging" produces nothing repeatable. Fourteen shots at 2.9 seconds is a specification. The flow recipes page shows the pinned pipelines built around specific formats.
Where it fails, measured
One flow in five fails, so it is worth knowing which one. It also means a failed flow is not a reason to change plan, since the credit cost is not what failed.

| failure | count | share |
|---|---|---|
| analysis failed on the source | 17 | 61 percent |
| analysis stopped unexpectedly | 7 | 25 percent |
| rejected by content moderation | 3 | 11 percent |
| video longer than the 120 second limit | 1 | 4 percent |
That matters because it changes what you do when it fails. Regenerating, switching engine or rewriting your prompt does nothing for a failure that happened at the analysis stage. The model list is not the lever here, and neither is lip sync quality.
The single most useful finding: paste less, upload more
Broken out by where the source came from: Platforms are explicit that automated downloading is against their terms, so use a video you have the right to use: YouTube's terms ↗ and TikTok's terms ↗ both cover this, and what you may publish is a separate question worth settling first.

| source | flows | average length | failure rate |
|---|---|---|---|
| Instagram link | 48 | 32.0s | 22.9 percent |
| TikTok link | 39 | 38.4s | 17.9 percent |
| YouTube link | 35 | 35.0s | 28.6 percent |
| uploaded file | 13 | 28.9s | 0 percent |
Thirteen uploads, zero failures. Every link source fails between roughly one in six and one in three.
The reason is straightforward once you see the failure table: pasting a link means something has to fetch that video first, and platforms actively make that unreliable. Uploading skips the step that breaks. If a link fails, do not retry the link. Download the file and upload it. That one change removes the largest failure class we have.
It is also worth knowing that YouTube links fail most, at 28.6 percent, despite YouTube being the easiest platform to get video from in principle. Longer, more varied source material is harder to segment cleanly than a phone shot vertical clip.
The 120 second ceiling
Sources are capped at two minutes. One flow in the dataset hit it, so it is not a common wall, but it is a hard one. If your source is genuinely longer than two minutes, cutting it down first is better than fighting the cap, and ffmpeg ↗ will trim without re-encoding.
The cap is not arbitrary. At 2.9 seconds a shot, a two minute source is already about forty clips, and each clip is a separate generation. Cost and failure surface both scale with shot count, which is the real reason short sources work better here. A 30 second source is the sweet spot and, conveniently, it is also the length that performs on TikTok and Instagram.
Before you paste anything, check these
Five checks, all free, that between them avoid most of the 20.7 percent. Running these first is also how you tell a source problem from a product problem, which is the same distinction that decides every lip sync failure.
- Length under 120 seconds, and ideally under 40. Shorter sources segment more cleanly and cost less.
- Have the file, not just the link. If you can download it, upload it. Zero of thirteen uploads failed.
- Count the cuts. If the source cuts every three seconds or faster, it will replicate well. If it is one long take, there is nothing to segment and you want a talking avatar instead.
- Check the content will pass moderation. Three flows were rejected at this step, and it happens before you are charged.
- Confirm the format is actually repeatable. A video whose appeal is one specific person doing one specific thing does not become yours by copying its cuts.
What this does not do
Being clear about limits is more useful than another feature list. It also does not clone a voice. That is its own pipeline with its own consent questions, covered in voice cloning and copyright.
Where this fits against the alternatives
Against editing by hand, the win is the cut list. Finding fourteen shot boundaries in a 34 second video and matching them is an hour of careful work, and the pipeline does it as the first step. What you still own is taste. And against doing nothing, the win is simply that a format you can specify is a format you can repeat, which is the whole argument for working from flows rather than from inspiration.
Against a text prompt, the win is repeatability. A prompt gives you something in the neighbourhood; a shot list gives you the same structure every time, which is what a format is.
Against stock template tools, the win is that you choose the reference. Templates are somebody else's formats, already saturated by the time they ship. A video you found this morning is not.
If you would rather start from a format that already works rather than find your own, the flows are pinned pipelines built exactly that way, and the avatar tools comparison covers the products around them.
Doing it, in order
Find a video under 40 seconds that cuts fast. Download the file rather than copying the link. Upload it, let the analysis produce the shot list, and check the shot count looks like the original. Regenerate the shots you want changed and leave the rest. Add your own audio. If the result is worth keeping, the publishing rules per platform decide how it has to be labelled.
The whole thing costs about 4.8 credits on average, which is roughly what one careless retry of a longer render costs, and you can try a flow ↗ before deciding whether the format is worth your time.
Ready to Create Your Own AI Avatar?
Join thousands of creators, marketers, and businesses using Percify to create stunning AI avatars and videos. Start your free trial today!
Get Started FreeGot questions?
Frequently asked
The source is analysed to find every cut, split at those boundaries, and each shot is regenerated as your version before the pieces are reassembled in order. Across 135 real replications the average source was 34.2 seconds and became about 14 clips of roughly 2.9 seconds each, for about 4.8 credits.
Because of the source, not the models. Of the failures we recorded, 61 percent were the analysis failing on the source video and another 25 percent were analysis stopping unexpectedly, so nearly nine in ten failures happen before any clip is generated. Regenerating or changing engine does nothing for those.
Upload the file. Across 135 flows, uploaded files failed 0 times out of 13, while Instagram links failed 22.9 percent of the time, TikTok links 17.9 percent and YouTube links 28.6 percent. Fetching video from a platform is the single largest failure class, and uploading skips it entirely.
Measured across 843 generated clips, the average shot runs 2.9 seconds, so a 34 second video carries roughly 14 separate shots. Videos that cut slower than about five seconds a shot read as talking videos rather than shortform, which is usually why a hand made version feels slow next to its reference.
120 seconds. The cap exists because shot count scales with length: at 2.9 seconds a shot a two minute source is already around 40 separate generations, and both cost and failure surface scale with that. Sources under 40 seconds work best.
No. You get the structure and the shot rhythm, not the performer. Putting a specific face in the result is a separate step handled by avatar generation, and audio is yours to supply along with its rights.
