Write a line, get a video

Describe the shot you want and a model makes it. Thirty four text to video models in one place, from 3 credits, no watermark on any plan, and you own what you make.

Make a video from text

Real text to video output

Each of these was generated from a written description by the model named under it. Nothing was uploaded.

How it works

1. Describe the shot

One line covering what is in frame, what moves, and what the camera does. Specific beats poetic.

2. Pick a model

Try the idea on a cheap model first. The playground quotes the exact cost before anything runs.

3. Generate and download

Run the take you want on a stronger model and download the MP4. No watermark on any plan.

The models

ModelFromWhat it is good at
Pixverse V63 creditsThe cheapest way to see whether an idea works at all
P-Video4 creditsCheap and quick for a rough pass
Wan 2.2 Ultra-Fast 480p10 creditsFast turnaround when you are iterating
Ovi15 creditsGenerates its own audio alongside the picture
Pika 2.220 creditsClean, well composed shots
Kling Video O3 Standard26 creditsSteady motion at a middling price
Seedance 2.0 Mini30 creditsHolds detail through movement
Kling V3 Turbo Standard34 creditsPeople who stay themselves as they move
Hailuo 2.334 creditsExpressive movement and camera work
LTX-2 Pro36 creditsLonger takes without a jump in price
Gemini Omni Flash Video39 creditsGoogle, for prompts that need reading carefully
Seedance V1.5 Pro42 creditsNatural motion on a bigger budget

Twenty two more text to video models are live beyond this table, thirty four in all, out of 130 models on the platform. Models differ in how they charge: some bill per second and some a flat rate for a fixed length, so the figure above is where each one starts rather than the price of a finished clip. The playground quotes the exact cost before anything runs. Plan prices are on the pricing page. Checked 23 September 2026.

Text to video, image to video, or a talking photo

These are three different jobs and it is worth knowing which one you have before you spend a credit.

  • Text to video is this page. You have nothing but an idea, and the model invents the whole shot from your description.
  • Image to video is for when you already have the picture and want it to move. Start at image to video.
  • A talking photo is a different family again: lip sync driven by an audio track, where the mouth and jaw follow a voice. Start at make a photo talk.

All three run on the same credits and the same API key, so moving between them is a parameter rather than a migration. API pricing.

Frequently asked questions

Everything you need to know before you start.

You write a line describing a shot and a model generates the video from it. Nothing is uploaded: the description is the only input, which is the difference from image to video, where you start from a picture you already have.

It depends on the model, and the playground quotes the exact cost before anything runs. The cheapest text to video model on Percify starts at 3 credits. Plan prices are on the pricing page.

Thirty four text to video models are active and public today, out of 130 models in total. They all run on the same credits and the same API key, so changing model is a parameter rather than a migration.

No. No plan applies a watermark, and under the terms you own the content you create through the service.

Not from text alone. Lip sync is a different family of models driven by an audio track, where the mouth and jaw follow the voice. If you want a photo that talks, start at make a photo talk.

Then you want image to video, which animates a still you upload rather than generating the whole shot from a description. It runs on the same credits.

It varies by model, and each one states its own duration in the playground before you run it. Some charge per second and some charge a flat rate for a fixed length, which is why the playground quotes the cost rather than this page.

Make a video from a line of text

Thirty four models, one set of credits, and the exact cost quoted before anything runs.

Make a video from text