Glossary
Video generation
Video generation produces moving footage from a written description or a still image, and it works in short clips, so anything longer is assembled from pieces rather than generated whole.
In plain terms
Describing a shot and getting moving pictures back. It is genuinely good at a few seconds at a time. It is much weaker at making the next few seconds look like they belong to the same thing, which is why almost everything longer is stitched together by a person afterwards.
Why it matters
Because the demonstrations are clips and the work is rarely a clip. A single striking shot is what gets shared and what a tool is judged on, while the actual job is usually a sequence where the same person, product and lighting have to survive from one shot to the next. That second thing is the hard part and it is where the tools differ most.
How it works
Output arrives in short pieces, and the length limit is structural rather than a setting. Models generate a bounded run of frames at a time, so anything of real duration is several generations assembled afterwards. The editing step is not a shortcoming of a particular product; it is the shape of the whole category.
Continuity across those pieces is the real problem, and it is what the current generation is competing on. Holding one character, one setting and one look steady from shot to shot is far harder than making any single shot beautiful, and the tools that lead on narrative work are the ones that carry a subject forward rather than the ones with the highest single-clip ceiling. Those are frequently not the same product.
Sound has moved from an afterthought into the generation itself. Several models now produce dialogue, ambient noise and effects together with the pictures rather than leaving them to be laid on later, which removes a whole editing stage and also means the audio is now something to art-direct rather than to add.
Control and spectacle are separate axes, and a tool can lead on one while being ordinary at the other. Directing means camera moves you asked for, references that hold a style, and editing that extends footage you already have, none of which one impressive clip demonstrates. The leading vendors now market themselves on both at once, which makes a demonstration a weaker guide than it used to be.
Cost and licence are settled before the creative question in commercial use. Generation is metered, longer and higher-resolution output costs more, and what you may do with the result depends on the training data behind the model and the terms attached to it. A tool trained on licensed material with indemnification available is a different commercial proposition from one that is not.
The two things people mean by good
Seen in the wild
A tool built for directed work, with real camera control, motion painting and an editing suite that extends and restyles existing footage.
RunwayA model whose pitch is narrative continuity: carrying a character and a voice from one shot into the next rather than producing one striking clip alone.
KlingA flagship model producing short cinematic clips with synchronised audio arriving alongside the pictures rather than added afterwards.
Veo (Google)
Common misconceptions
People assume
It can make a finished video.
In fact
It makes shots, and somebody assembles them. Length is bounded per generation, so a sequence of any duration is several outputs joined afterwards, and the joining is where continuity problems become visible. Budget for an edit, always.
People assume
The most impressive demo is the best tool.
In fact
Demonstration clips test the axis that matters least for commissioned work. Single-shot spectacle and directability are different strengths, and a clip can only ever show the first however loudly a vendor markets the second, so what a demonstration cannot answer is what happens on the second shot.
People assume
The output is yours to use commercially.
In fact
That depends on the model's training data and the terms attached to it, not on the fact that you generated it. Some products are built explicitly for commercial use with indemnification available; others are not, and the difference matters long before anybody discusses quality.
Telling them apart
Video generation vs Text to video
Video generation
The whole capability, however it is prompted.
One route in: a written description, nothing else.
Most tools also take a still image as the starting frame, which usually gives far more control than words alone, so the narrower phrase describes only part of what is on offer.
Questions
- How long can a generated clip be?
- Short, and the limit is a property of the category rather than of any one product. Models produce a bounded run of frames per generation, so anything of real duration is assembled from several. Treat the tool as a camera that shoots takes rather than as something that delivers a finished sequence.
- Why does my character keep changing between shots?
- Because each generation is a fresh act of creation unless something carries the subject forward. Reference images, character-consistency features and multi-shot support all exist for exactly this, and they are the axis worth comparing when the work is a sequence. A tool with no answer here will be beautiful and unusable for narrative.
- Does it produce sound as well?
- Increasingly yes, generated together with the pictures rather than added afterwards. Dialogue, ambience and effects arriving in sync removes an editing stage, and it also means the audio becomes something to direct rather than something to source. Not every model does this, so it is worth checking rather than assuming.
- What decides whether we can use the output commercially?
- The model's training data and the terms attached to it, which is a procurement question rather than a creative one. Products built for commercial use train on licensed material and may offer indemnification; others make no such commitment. Establish this before a campaign is built on the output rather than after.
Key takeaways
- It generates shots, not sequences; assembly is always somebody's job.
- Continuity across shots is the hard problem and the real differentiator.
- Spectacle and directability are separate axes, and rarely the same tool.
- Audio now arrives with the pictures, which removes a stage and adds a decision.
- Commercial use depends on training data and terms, not on who generated it.
Tools that use this
- Runway
Directed work: camera control, motion painting, editing existing footage.
- Kling
Narrative continuity, carrying a character across shots.
- Veo (Google)
Short cinematic clips with audio synchronised at generation.
Last checked July 2026