Quick Answer
The AI avatar era runs from a lip sync research paper in 2020 to open, audio driven talking video in 2025 and 2026. The dated landmarks: Wav2Lip at ACM Multimedia 2020; SadTalker at CVPR 2023, which animates one photo; OpenVoice powering instant voice cloning from May 2023 and MIT licensed in April 2024; Meta's AI labels through 2024; LatentSync 1.5 and 1.6 in 2025; InfiniteTalk's open weights on 19 August 2025; and LongCat Video Avatar 1.5 in May 2026. Today the work is a digital twin of you, priced by the second and labeled on the platforms.
A dated AI avatar timeline, checked on 15 September 2026: Wav2Lip in 2020, SadTalker in 2023, Meta's AI labels in 2024, LatentSync and InfiniteTalk in 2025, LongCat in 2026, and what it costs now.
Related next steps
Most AI avatar timelines are written from memory. This one gives a date and a source for every entry, taken from the research repositories and the platforms' own announcements, read on 15 September 2026. Where no primary date exists, the entry is left out.
2020 to 2022: lip sync becomes a research problem people can run
- 2020: Wav2Lip ↗, presented at ACM Multimedia, syncs a face in a video to any speech. Its open version is non commercial, and the repository points to a commercial model instead. Every "make the mouth match the audio" tool of the next years traces back to it.
What existed then: research code that needed a GPU and a video to work from, not a product.
2023: one photo is enough, and voices clone in seconds
- 2023: SadTalker ↗, at CVPR, animates a single photo from audio. Its README later records the licence moving to Apache 2.0, with the non commercial restriction removed.
- May 2023: OpenVoice ↗ starts powering the instant voice cloning in MyShell, by its authors' own account.
What changed: you no longer needed video of a person, only a photo, and a voice could be copied from a short sample.
2024: platforms write the rules, and licences open up
- 6 February 2024: Meta says it will require ↗ people to disclose organic posts with photorealistic video or realistic sounding audio that was digitally created or altered, and that it may apply penalties.
- April 2024: OpenVoice ↗ is released under MIT, which its authors describe as free for commercial use; 16 April 2024, MuseTalk ↗ ships a public demo on Hugging Face.
- 1 July 2024: Meta renames the label ↗ people see from "Made with AI" to "AI info".
What changed: the question stopped being whether the video could be made and became whether it was labeled.
2025: audio driven video, open and long
- 14 March and 11 June 2025: LatentSync ↗ 1.5 and 1.6 from ByteDance, lip sync built on latent diffusion, with 1.6 trained at higher resolution.
- 19 August 2025: InfiniteTalk ↗ from MeiGen releases its technique report, weights and code under Apache 2.0: talking video of unlimited length from a photo or a video, with head, posture and expression following the audio.
- 16 December 2025: MeiGen releases LongCat Video Avatar, a unified model for audio driven human video.
What changed: the models behind hosted avatar tools became downloadable, so tools compete on price, workflow and rights.
2026: your own twin, priced by the second
- 21 May 2026: MeiGen releases LongCat Video Avatar 1.5, which it calls an upgraded open source framework for audio driven human video.
- Now: the tools sell a twin of you rather than a stock presenter. On Percify, Clone Yourself makes your face and voice for 5 credits each, once, and talking video ↗ costs 2 credits a second. TikTok requires a label ↗ on realistic AI generated video, and Meta requires disclosure.
What it means: the hard part is no longer the model. It is the photo, the voice sample, the licence and the label.
Where it landed: the tools on 15 September 2026
| Tool | Free tier | An avatar of you | Watermark |
|---|---|---|---|
| Percify | 10 credits, no card | Your face and voice, 5 credits each, once | None on any plan |
| HeyGen | 3 videos a month, up to a minute | 1 custom avatar on the free plan, 500+ stock | Removal from Creator, $29 a month |
| Synthesia | Basic: 1,200 credits a month, no downloads | 5 personal avatars from Creator, €79 or €58 billed yearly | Logo removed from Starter, €26 or €16 billed yearly |
| D-ID | 14 day trial, 3 minutes, personal use | Any photo talks, with no avatar step | Full screen watermark on the trial |
| Vidnoz | 30 credits a day, about 60 seconds | 1,900+ stock avatars listed | 720p with a watermark on the free plan |
Each row was read on that maker's own pricing page on 15 September 2026, and we make Percify, so read its row with that in mind. The pattern across six years: the model stopped being the product. What is sold now is whose face it is, what it costs a second, and whether you may publish the file.
See the 2026 trends, with sources, free voice cloning compared, free lip sync generators and the best AI avatar generators compared.
Ready to Create Your Own AI Avatar?
Make a talking video of yourself from one photo, in your own voice. No watermark on any plan.
Get Started FreeFree to start · no credit card
Got questions?
Frequently asked
As research anyone could run, 2020: Wav2Lip was presented at ACM Multimedia that year and syncs a face in a video to any speech. Animating a single photo followed with SadTalker at CVPR 2023.
The models opened up. ByteDance's LatentSync reached 1.5 in March 2025 and 1.6 in June, and MeiGen released InfiniteTalk's weights and code under Apache 2.0 on 19 August 2025, making talking video of unlimited length from a photo or a video.
Meta said on 6 February 2024 that it would require disclosure for photorealistic video or realistic sounding audio that was digitally created or altered, and on 1 July 2024 it renamed the viewer label from "Made with AI" to "AI info". TikTok requires creators to label realistic AI generated content, including AI generated speech.
On 15 September 2026, Percify charged 5 credits each for your face and voice, once, and 2 credits a second of talking video, so a minute costs 120 credits. Other tools sell monthly plans with stock presenters instead.
Read on 15 September 2026: Percify gives 10 credits with no card and adds no watermark on any plan; HeyGen's free plan allows 3 videos a month of up to a minute; Synthesia's Basic plan gives 1,200 credits a month but no downloads; Vidnoz gives 30 credits a day, about 60 seconds, at 720p with a watermark; D-ID offers a 14 day trial of 3 minutes with a full screen watermark, for personal use.
Several are. InfiniteTalk and LatentSync are Apache 2.0, SadTalker moved to Apache 2.0, and OpenVoice is MIT and free for commercial use. Wav2Lip's open version is non commercial, and some voice models release code openly while their weights carry stricter licences.
