AI Training Video · 2026 Guide

AI Training Videos: How to Make Them

An AI training video is a training module whose on-screen presenter is generated rather than filmed. You write a script, supply one photo, choose a voice, and the model produces a presenter who delivers it with matched lip movements. No studio, no camera, no scheduling — and when the policy changes you edit the script and regenerate instead of re-filming.

That last point is the one that matters in practice. Filmed training decays: re-recording means rebooking a presenter, a room and an editor, so revisions get postponed and staff end up trained on material that is out of date. Removing the reshoot cost is what makes a curriculum maintainable.

How to make an AI training video, step by step

  1. 1

    Write the script first

    Write the module as spoken words, not slide bullets. One idea per paragraph, and read it aloud once — the script is the only real production constraint left, so it is where the quality is decided.

  2. 2

    Choose or generate the presenter

    Upload one clear, front-facing photo to generate a presenter, or reuse an existing one so the whole curriculum keeps a consistent face. Using a colleague’s likeness requires their consent.

  3. 3

    Pick the voice

    Choose from the stock voice models, or clone a voice with the speaker’s consent so every module in the curriculum sounds like the same trainer.

  4. 4

    Generate the video

    The model synthesises the narration and regenerates the presenter’s lip movements to match it. Generation runs in the background, so you can queue the next module while it works.

  5. 5

    Review, then localise if needed

    Check the opening and closing seconds, then download. To deliver the same module in another language, dub it — the lip movements are regenerated for the new audio rather than voiced over.

AI presenter vs filming vs slides with voice-over

Most L&D teams are choosing between three options, and they fail in different ways. Slides with a voice-over are cheap but have no presenter, and completion rates suffer for it. Filming looks best on day one and is the hardest to keep current. A generated presenter sits between them: it keeps a face on screen and stays cheap to revise.

ApproachCost to updatePresenter on screen
Filmed presenterReshoot: rebook presenter, room, editorYes
Slides + voice-overLow — re-record narrationNo
Generated presenterEdit the script, regenerateYes

Onboarding videos

Onboarding is the highest-repetition content in the business: the same twenty minutes delivered to every new starter forever. It is also the content most often delivered live by a manager, which makes it inconsistent and expensive. Generating it once fixes both — every starter gets the same version, and the version can be corrected the day a process changes. See AI video for enterprise teams for how this fits alongside internal comms.

Compliance training

Compliance is where the reshoot problem bites hardest, because the material is legally required to be current and changes on someone else's schedule. Keeping the presenter, pacing and structure identical between revisions means a regulatory update changes only the wording that actually changed — which is also what makes the revision easy to review and sign off.

Product and skills training

Product training goes stale every release. Writing modules against a script that lives in version control alongside the release notes means the training can ship with the feature rather than months later. For customer-facing product training, paid plans include commercial rights and produce videos with no watermark.

Delivering the same course in other languages

A finished module can be dubbed into other languages with the presenter's lip movements regenerated to match the new audio, rather than a voice-over laid over unchanged footage. For a distributed workforce this replaces commissioning a separate shoot per region. The full process is in the AI video dubbing guide.

What it costs

Cost is per generation in credits rather than per seat or per studio day, so a curriculum costs roughly what it takes to render it. Lip-sync generation starts at 2 credits on the fast model and 4 on the standard one, failed runs are refunded automatically, and there is a free tier for testing a module before committing. Full numbers are on pricing.

Getting a training video that people finish

Frequently asked questions

What is an AI training video?

A training video where the on-screen presenter is generated rather than filmed. You supply a script and a photo, and the model produces a presenter who speaks it with matched lip movements. Nothing is filmed, so there is no studio, no camera and no reshoot when the content changes.

How long does it take to make one?

Minutes rather than days. The generation itself runs in the background while you work on the next module, so the real constraint is how quickly you can write the script — not scheduling a presenter, a studio or an edit suite.

What happens when a policy or product changes?

You edit the script and regenerate. This is the difference that matters for compliance and onboarding content: filmed training decays because re-recording means rebooking everyone, so teams postpone updates and staff are trained on stale material. A generated presenter has no rebooking cost.

Can the same course be delivered in other languages?

Yes. A finished video can be dubbed into other languages with the speaker’s lip movements regenerated to match the new audio, so it does not look like a voice-over pasted on top. See the AI video dubbing guide for the full process.

Do we own the videos and can we use them internally and commercially?

Yes. Paid plans carry commercial usage rights and produce videos with no watermark, which is what makes them usable for customer-facing product training as well as internal onboarding.

Do we need a real person to appear in the training?

Not necessarily. You can generate a presenter from a single photo, or use a consistent brand presenter across a whole curriculum. Using a real colleague’s likeness requires their consent, and voice cloning requires the speaker’s own consent.

What does it cost per video?

Cost is per generation in credits, not per seat or per finished minute of studio time. Lip-sync generation starts at 2 credits on the fast model and 4 on the standard one, and failed runs are refunded automatically. There is a free tier to test a module before committing.

Is this suitable for compliance training specifically?

It suits it well, because compliance content is the material that goes stale fastest and is the most expensive to re-film. The presenter, pacing and structure stay identical between versions, so a revision changes only the wording that actually changed.

Make your first training module free

One photo and a script. No studio, no camera, no reshoot when it changes.

Start free

More Percify guides