Category
Text to video generators, and the four different things they mean by it
Four categories of software are sold under one phrase, and they do not produce the same kind of picture, cost the same money, or fail in the same way. Below: the same paragraph put through all four, and what each one gives back.
Published 13 September 2026 · 10 min read · prices carried from this site’s vendor-page checks on 10 September 2026
The short version
- A person should say it: avatar tools. Synthesia or HeyGen, $29 a month. The picture is a presenter; the script is read aloud verbatim.
- You have writing and no pictures: stock-footage tools. Pictory or InVideo AI, $17–$29 a month. Fast, cheap, and the footage is about your topic rather than about your product.
- You want footage that never existed: generative video. Extraordinary at six seconds, still not a ninety-second explainer, and you cannot dictate what appears.
- The text describes a thing, not a feeling: motion graphics. The subject gets drawn — charts, steps, interfaces, numbers — and the timing is cut to the narration. This is the category people usually meant.
“Text to video” means four different products
The phrase describes an input, not an output, which is why it has ended up attached to four unrelated things. Every one of them takes writing at one end. What comes out at the other is the part nobody puts in the marketing copy.
| Category | What appears on screen | Typical entry price | Fails when |
|---|---|---|---|
| Avatar | A synthetic presenter reading your script | $29/mo (Synthesia, HeyGen) | Nobody needed a presenter |
| Stock footage | Library clips matched to your sentences | $17–$29/mo (Pictory, InVideo AI) | The subject is your own product |
| Generative | Invented footage, a few seconds at a time | Varies wildly; credit-metered | You need a specific thing shown |
| Motion graphics | The subject drawn: charts, steps, interfaces | €3 per video (ExplainVids) | You wanted a human face |
For the eleven-tool shortlist rather than the four-category map, the explainer video software comparison puts them in one table with published prices and a cost-per-video column.
The same paragraph, four ways
Take one ordinary piece of product writing — the kind actually sitting in a doc somewhere waiting to become a video:
There are three concrete things in that sentence: a four-step sequence, a duration, and a percentage. What each category does with them is the whole difference between the products.
| Category | What you get back | What happened to the 40% |
|---|---|---|
| Avatar | A presenter in a studio set, speaking the sentence. Competent, clear, and visually identical to every other avatar video. | Spoken aloud. Never shown. |
| Stock footage | Someone smiling at a laptop, a stock team high-fiving, a stock hand on a stock phone. Pleasant, generic, about nothing. | Spoken aloud, possibly captioned over unrelated footage. |
| Generative | Six seconds of striking, slightly dreamlike footage that is not your product and cannot be made to be. | Gone. The model does not draw figures. |
| Motion graphics | Four numbered steps appearing in sequence, a timer, and a bar falling from 100 to 60 as the words are spoken. | Drawn as a falling bar, on the syllable it is said. |
This is the test worth running before paying for anything. Write one paragraph containing a number you care about, put it through a free tier of each category, and see which one puts the number on screen. Three of the four will say it and not show it.
Avatar tools: a presenter reading your text
Synthesia and HeyGen are the category, both at $29 a month to start. You type a script, pick a presenter, and get a person on screen delivering it, with lip sync good enough that the average viewer does not question it and dubbing into dozens of languages that genuinely is the best in the market.
Right when
- The video is training, compliance or an internal announcement, where a person addressing you is the correct form.
- The same script has to exist in eight languages with matched lips. Nothing else does this as well.
- You are replacing a video that would otherwise be a real person on a webcam.
Wrong when
- The subject is a product, a process or a comparison. The presenter describes it while the screen shows a presenter.
- It is going on a homepage. Avatar videos are now recognisable as avatar videos, and that recognition happens before your first claim lands.
If you are already paying for one of these and the videos feel competent but flat, Synthesia alternatives grouped by why you are leaving covers the four exits, including the one that leaves the category altogether.
Stock-footage tools: your text over other people’s clips
Pictory and InVideo AI, $17–$29 a month. These split your writing into sentences, search a stock library for something plausibly related to each, and lay the clips end to end under a synthetic voice. For turning a blog post into something postable, they are genuinely the fastest route that exists.
The limit is structural rather than a quality problem: the pictures come from a library, so they can only ever be about your topic. A tool searching for “onboarding” returns a person at a desk. It cannot return your onboarding, because no stock library has footage of your product in it.
Generative video: invented footage, and the length problem
Sora, Runway and Veo invent footage frame by frame from a description. At their best the results are extraordinary, and they are the only category here that can show you something that has never been filmed.
Two things keep them out of explainer work for now. The first is length: these models produce clips measured in seconds, so a ninety-second explainer becomes an editing project stitching many generations together — which is the timeline you were trying to avoid. The second is control. You cannot tell a generative model to put your pricing table on screen, spell a product name correctly, or draw a bar chart that matches your actual figures. It produces something in the neighbourhood of the description, and the neighbourhood is not close enough when the video contains claims.
For a mood piece or a title sequence, excellent. For a video whose job is to explain a specific thing accurately, the category is not there yet, and a per-second credit meter makes finding that out expensive.
Motion graphics: the text drawn as the thing it describes
The fourth category takes the same script and, instead of finding someone to read it or something to illustrate it, draws the subject. A four-step process becomes four steps appearing in sequence. A percentage becomes a bar that moves. An interface becomes an interface.
This is the oldest form of explainer video and, until recently, the most expensive: it is what agencies charge $3,000 to $15,000 for a minute of, because historically it meant a designer in After Effects moving keyframes until the animation agreed with the voiceover. What changed is that the agreement between animation and voice — the part that ate the hours — can now be measured rather than nudged.
The trade is the mirror image of the avatar category: there is no human face, so if the point of the video was a person addressing the viewer, this is the wrong instrument.
What it costs per finished video
Monthly prices hide the number that matters. A $29 subscription is $29 a video if you make one that month and $2.42 a video if you make twelve — and almost nobody makes twelve.
| Videos a month | Synthesia $29 | HeyGen $29 | Vyond $99 | ExplainVids |
|---|---|---|---|---|
| 1 | $29.00 | $29.00 | $99.00 | €3.00 |
| 4 | $7.25 | $7.25 | $24.75 | €3.00 |
| 12 | $2.42 | $2.42 | $8.25 | €2.92 |
The honest reading of that table: at twelve videos a month a subscription wins on unit price, and at one or four it does not. Most people asking what a text to video generator costs are making one. The full breakdown of what an explainer video costs runs the same arithmetic against freelancers and agencies too.
Which one to use for which job
| What you are making | The category that fits |
|---|---|
| Compliance or safety training | Avatar. A person addressing the viewer is the form. |
| The same video in eight languages | Avatar. Dubbing is the category’s real advantage. |
| A blog post, on LinkedIn today | Stock footage. Fastest path from writing to postable. |
| A title sequence or a mood piece | Generative. The only one that invents imagery. |
| A product, a process, or figures | Motion graphics. The subject gets drawn. |
| A landing page hero | Motion graphics. The page is the subject — and a URL can be the whole input. |
Where ExplainVids fits, and where it does not
We are the fourth category and nothing else. You type what the video should explain — or paste a URL and we read the page — and about two minutes later there is a narrated film with motion graphics, a real voice, captions lit word by word, vertical or widescreen, for €3 with no subscription. Your figures are drawn as charts because you gave them to us, not invented because a model guessed.
Pick us when
- The text contains something specific — steps, numbers, a comparison, a workflow — that should appear on screen rather than merely be read out.
- You want one video without a subscription decision attached to it.
- The thing you want a video of already exists as a web page.
Pick something else when
- A human face is the point — a founder message, a piece to camera, a training module with a presenter. Use an avatar tool.
- You want a long library of near-identical videos in many languages. That is Synthesia’s job and it does it better.
- You want invented cinematic footage. That is generative video, and it is a different product entirely.
The fastest way to find out which category you actually needed is to put one paragraph with a real number in it through this and watch where the number ends up. Three euros, about two minutes.
Make one and seeQuestions
What is a text to video generator?
A tool that takes writing and returns a finished video without an editing timeline. The phrase covers four incompatible products: avatar tools that put a synthetic presenter on screen reading your words, stock-footage tools that match library clips to your sentences, generative models that invent new footage from a description, and motion graphics tools that draw the subject your text describes. All four are sold as “text to video” and only one of them will be what you meant.
Which text to video generator is best?
It depends entirely on what the picture is supposed to show. If a person addressing the viewer is the point, Synthesia or HeyGen at $29 a month. If you have a blog post and want it on LinkedIn by lunchtime, Pictory or InVideo AI at $17–$29 a month. If you want a few seconds of footage that never existed, that is generative video and a different category again. If the text describes a product, a process or a set of numbers, you want motion graphics — the subject drawn rather than a person describing it.
Is there a free text to video generator?
There are free tiers, and they share the same three limits: a watermark, a length cap, or a count cap — usually two of the three. Synthesia’s free Basic tier is about 10 minutes a month of credits, HeyGen’s is three videos a month at up to a minute each and watermarked, Colossyan’s Starter is free with 20 minutes a month. If the video is going on a homepage or into an ad, none of them apply, because the watermark is the product.
Can a text to video generator use my own script?
Every category accepts a script, but they do different things with it. Avatar tools read it out verbatim, so the script is the whole input. Stock tools chop it into sentences and search a library for each one, which means the writing controls the pictures only indirectly. Generative tools treat it as a prompt and largely rewrite what you asked for. Motion graphics tools sit in between: the script sets the narration and the narration sets the timing, so a rewritten line moves everything after it.
What is the difference between text to video and AI video generation?
“AI video generation” usually means the generative category specifically — a model inventing footage frame by frame from a description, as Sora, Runway and Veo do. “Text to video” is the older and broader phrase and includes the three assembly-based categories, which do not invent footage at all: they perform it, retrieve it, or draw it. The distinction matters because the generative category is the only one where you cannot control what appears.
How long does a text to video generator take?
Minutes, in every category, which is the whole point of the format. The honest number is not the render time but the number of attempts: a tool that returns a video in ninety seconds but needs six goes to get the pictures right has cost you an afternoon. Ask how much of the output is determined by your input, because that is what decides how many attempts it takes.
How these prices were checked
No price on this page was re-read for it. Every figure is carried from the two articles on this site that did read them off the vendor’s own pricing page — the software comparison on 6 September 2026 and the Synthesia piece on 10 September 2026. Restating a number from memory is how one site ends up quoting two different prices for one plan, so the figures have exactly one home and this page borrows them.
Generative video is deliberately left without a price. Those tools meter by credit against clip length and resolution, and any single monthly figure would misrepresent them. These plans change often: if you are reading well after 13 September 2026, check the one you are about to buy.
Related: explainer video styles once you know it is motion graphics you want, corporate explainer videos if this is for training or internal comms, fifteen examples worth studying for what good looks like, and the explainer video generator itself if you would rather just try one.