Glossary category
Modalities and media
The kinds of input and output AI works with.
16 terms
Able to work with more than one kind of input or output: text, images, audio, sometimes video.
The field concerned with getting software to interpret what is in an image or video: objects, faces, defects, movement.
Producing a new picture from a written description rather than retrieving or editing an existing one.
The kind of model behind most image generators, which starts from visual noise and refines it step by step towards the picture your description asks for.
AI you speak to and that speaks back, joining speech recognition, a language model and synthetic speech into one conversation.
Turning recorded or live speech into written words, which is what produces a meeting transcript or live captions.
Turning written words into spoken audio, used for narration, accessibility and assistants that answer out loud.
Reading the text inside a picture of a page, so a scan or a photograph becomes words a computer can search and copy.
- Vision language model (VLM)
A model that reads images and text together, so it can answer questions about a screenshot, chart, form or photograph.
- Document understanding
Pulling structure and meaning out of documents such as invoices or contracts, rather than only converting them to plain text.
- Video generation
Producing moving footage from a written description or a still image, typically in short clips that are then edited together.
- Image editing
Changing an existing image by instruction, such as removing a background or replacing an object, rather than generating one from nothing.
- Text to video
Turning a written description straight into a video clip, the video counterpart of asking for an image in words.
- Voice cloning
Building a synthetic copy of a particular person's voice from samples, which raises a consent question before any practical one.
- Avatar video
A presenter-style video built from a synthetic person reading a script, used for training and internal communications without filming anyone.
- Audio generation
Producing music, sound effects or other audio from a written description, distinct from reading text aloud in a chosen voice.
Last checked July 2026