Skip to content

Glossary

Text to video

Text to video turns a written description straight into a clip, which is the fastest route in and the least precise, because words specify a shot far more loosely than a picture does.

In plain terms

Describing a shot and getting moving pictures back, with nothing to work from but the sentence. It is the quickest way to try the idea and the vaguest way to ask for anything specific, because a sentence leaves almost everything undecided. Most people who use this seriously stop starting from words fairly quickly.

01

Why it matters

Because the phrase names the route people try first and abandon soonest. A description leaves the framing, the light, the palette and the exact subject to be invented, so the result is frequently striking and rarely the thing that was wanted, and the fix is not a better sentence but a different starting point.

02

How it works

A sentence underspecifies a shot by an enormous margin, which is the whole difficulty. Every element a director would decide is left open, so two runs of the same description produce two legitimate interpretations, and neither is wrong. What feels like unreliability is usually the request having been ambiguous in ways prose cannot easily fix.

Starting from a picture removes most of that ambiguity at once. Nearly every serious tool accepts a still image as the opening frame, which fixes the subject, the composition and the palette before any motion is decided, and leaves the description to do the one job it is good at: saying what should happen next. That combination is what most production use actually looks like.

Sound has moved into the same request in several products. Dialogue, ambience and effects now arrive generated alongside the pictures rather than being laid on afterwards, which removes an editing stage and quietly adds another thing the sentence is being asked to specify. More is being decided by the same few words.

The phrase describes an entry point rather than a category, and that trips buyers up. Products marketed under it also take images, references and camera instructions, so comparing them on the quality of their word-only output tests the route almost nobody uses in earnest. The useful comparison is what else they accept as a starting point.

What each starting point actually decides

What each starting point actually decidesThe reason the word-only route disappoints so reliably is that people compare it against a mental image they never wrote down. Somebody imagines a specific shot, describes it in a sentence, and receives something that satisfies every word they typed while matching almost nothing they pictured. The gap is not the system failing to understand; it is a sentence carrying perhaps a tenth of what was in their head, with the rest supplied by an interpretation that had no way of knowing. Handing over a still image collapses that gap immediately, because the picture contains all the decisions the sentence could not hold, and the words are then free to do the thing prose genuinely does well, which is to describe change over time. That reframing is worth more than any amount of practice at writing longer descriptions.A sentenceSubject, loosely.Mood, sometimes.Framing, light, palette: open.Two runs, two readings.A first frameSubject, exactly.Composition, fixed.Palette and light, settled.Words left to say what happens.The right-hand column is not aworkaround; it is what thecapability is for. The left-handone is how everybody tries itand how almost nobody keepsworking.
The reason the word-only route disappoints so reliably is that people compare it against a mental image they never wrote down. Somebody imagines a specific shot, describes it in a sentence, and receives something that satisfies every word they typed while matching almost nothing they pictured. The gap is not the system failing to understand; it is a sentence carrying perhaps a tenth of what was in their head, with the rest supplied by an interpretation that had no way of knowing. Handing over a still image collapses that gap immediately, because the picture contains all the decisions the sentence could not hold, and the words are then free to do the thing prose genuinely does well, which is to describe change over time. That reframing is worth more than any amount of practice at writing longer descriptions.
03

Seen in the wild

  • A model in this shape with native audio, whose distinguishing pitch is holding a scene, a character and a voice together across several shots.

    Kling
  • A flagship model generating short cinematic clips from prompts and reference frames, with dialogue and effects arriving alongside the pictures.

    Veo (Google)
  • A fast, playful tool taking both text and an image, with a family of effects and lip-sync tuned for output that reads instantly on a feed.

    Pika
04

Common misconceptions

People assume

A better description fixes an unpredictable result.

In fact

Beyond a point it cannot, because prose runs out of precision before a shot runs out of decisions. Framing, palette and the exact subject stay open however carefully the sentence is written, which is why the reliable move is supplying a first frame rather than rewriting the request again.

People assume

It describes a category of product.

In fact

It describes one way in, and the products under the label nearly all accept images, references and camera direction too. Comparing them on word-only output measures the route serious users abandon first, so it is a poor basis for choosing between them.

05

Telling them apart

Text to video vs Video generation

Text to video

One route in: a written description alone.

Video generation

The whole capability, however it is prompted.

The narrower phrase names the weakest input, which is why a tool described this way is usually better than the description suggests.

06

Questions

Why do two runs of the same description differ so much?
Because the description left most of the shot undecided and both results are legitimate readings of it. Words fix the subject loosely and leave framing, light and palette open, so variation is the system filling gaps rather than behaving inconsistently. Supplying a first frame closes most of those gaps in one move.
Is starting from an image cheating?
It is what most production use looks like. A still fixes composition, subject and palette before motion is considered, which leaves the description doing the job prose is actually good at, namely saying what should happen. Treating the image as the specification and the words as the direction is the working pattern.
Does the audio come from the same request?
In several products now, yes, generated alongside the pictures rather than added afterwards. That removes an editing stage and means the same short description is being asked to specify sound as well as vision, so it is worth stating explicitly what should be heard rather than leaving it to be inferred.
07

Key takeaways

  • A sentence underspecifies a shot; the variation is the gap being filled.
  • A first frame fixes subject, framing and palette in one move.
  • Audio increasingly arrives in the same request, adding to what words must carry.
  • The phrase names an entry point, so it undersells the products it labels.
09

Tools that use this

  • Kling

    Native audio and continuity of scene, character and voice across shots.

  • Veo (Google)

    Cinematic clips from prompts and reference frames, audio included.

  • Pika

    Text and image in, effects and lip-sync tuned for social output.

Last checked July 2026

All glossary terms