Glossary
Multimodal
Multimodal describes a system that works with more than one kind of material, taking in or producing some combination of text, images, audio and video rather than text alone.
In plain terms
It means you can show it a thing instead of describing the thing. Point a camera at a broken part and ask what it is. Paste a screenshot of a spreadsheet rather than retyping the figures. Hand it a recording rather than a transcript. The word covers an enormous range of actual ability, though, and being willing to accept an image is not at all the same as being good at reading one.
Why it matters
Most of the friction in office work is transcription: copying a figure off a scan, describing a layout in words, typing up what somebody said. Removing that step is where the practical saving lives, and it is larger than it sounds because the transcription step is also where errors enter. The catch is that supporting a kind of material tells you nothing about how well it handles it, and the gap between the two is wide enough that this is worth testing rather than believing.
How it works
Different kinds of material are converted into a common internal form so that one system can work across them. The mechanics differ between products and the consequence is the same: a photograph and a sentence about the photograph end up in a shape one model can handle. That is what lets you ask a question in words about something you supplied as a picture.
Taking material in and producing it are separate questions, and vendors rarely separate them for you. A model that accepts images may only ever reply in text, and generating images is usually a different model doing a different job. Read a capability claim as two claims, one about input and one about output, and check which one is being made.
Competence varies enormously within a single kind of material. Describing what is in a photograph, reading printed text inside it, reading values off the axes of a chart, counting objects and judging exact positions are different skills. A system that handles the first two comfortably can be unreliable at the last two, and it will not tell you which one you have asked for.
Everything consumes the same fixed budget. An image is worth a considerable number of tokens, so a handful of scanned pages fills a window far faster than the same content as text would. Where a workflow attaches images routinely, this is the constraint that bites first, and it usually arrives as an unexpected cost rather than as an error.
Some products chain separate models rather than running one that handles everything. Speech is transcribed by one and the text is passed to another, or an image is described before a language model sees the description. That works well and fails differently: detail is lost at the handoff, and any mistake in the first stage is inherited with confidence by the second.
Four things you might ask about one photograph
Seen in the wild
Photograph a whiteboard after a meeting and ask an assistant to turn what is written on it into a list of actions.
Google GeminiPaste a screenshot of a chart and ask what the trend is, then check the two or three values you can read yourself against what it told you.
ChatGPTAttach a scanned page and ask about a figure buried in the middle of it, which tests reading and locating at the same time.
Claude
Common misconceptions
People assume
Supports images means it can read our documents.
In fact
It means images will be accepted. Whether a particular scan, form layout, handwriting style or chart is read accurately is a separate question with a different answer for every product and every document type. Ten minutes with twenty of your own worst pages settles it far better than any specification sheet.
People assume
It sees the picture the way I do.
In fact
It handles some aspects well and others poorly, and gives no sign of which. Describing a scene is generally dependable. Counting items, judging which of two things is nearer, and reading small print in a corner are considerably less so, and the answer arrives in the same assured tone either way.
People assume
One model is doing all of it.
In fact
Often several are chained, with one transcribing or describing and another reasoning about the result. That is a perfectly reasonable design and worth knowing about, because it tells you where to look when something is wrong: usually at the first stage, whose mistakes the second stage repeats without hesitation.
Telling them apart
Multimodal vs OCR
Multimodal
Handles the image as a whole and can answer questions about it, including things nobody wrote down, such as what the layout implies.
Turns the characters in an image into text, precisely, and does not attempt to understand any of it.
For a clean page of printed text, the older technique is usually faster, cheaper and more accurate. For a photograph of a damaged part, it has nothing to offer at all.
Multimodal vs Computer vision
Multimodal
A general system that happens to accept images alongside text, answering in language about whatever you supply.
Systems built specifically for images, often trained for one task such as detecting defects or counting stock.
The general one is far easier to try and far harder to rely on at volume. Where a single visual task runs thousands of times a day, the purpose-built option is usually both cheaper and more accurate.
Questions
- Does multimodal mean it can create images as well?
- Not necessarily, and usually not with the same model. Accepting images and generating them are separate capabilities, and plenty of assistants do the first while handing the second to a different system entirely. When a vendor uses the word, ask which direction they mean, because the two are frequently conflated in the same sentence.
- How good is it at reading charts and tables?
- Better at the gist than at the figures. Describing the shape of a trend is usually reliable; reading precise values off axes, and getting every cell of a table across correctly, is much less so. Where the numbers matter, treat what comes back as a draft to check rather than as an extraction you can act on.
- Does sending images cost more?
- Considerably, in the sense that an image occupies far more of the budget than a paragraph does. A workflow that attaches scans routinely will find the fixed space filling much faster than expected. Where only part of a page matters, cropping before sending is a cheap and surprisingly effective habit.
- How do I decide whether it is good enough for our documents?
- Collect twenty pages that represent the hard cases rather than the typical ones: the faint scan, the unusual layout, the handwritten annotation, the rotated photograph. Run all twenty, check every answer against the page, and count. That number tells you more than any vendor claim, and it takes an afternoon.
Key takeaways
- The word means more than one kind of material, and says nothing about how well any of them is handled.
- Accepting a kind of material and producing it are separate claims, often conflated in one sentence.
- Ability varies sharply within a single kind: describing a photograph is easier than counting what is in it.
- Images consume far more of the fixed budget than the same content as text.
- Several models are often chained, so first-stage mistakes are inherited with confidence by the second.
Tools that use this
- Google Gemini
Accepts photographs alongside a written question about them.
- ChatGPT
A quick way to test a screenshot against values you can read yourself.
- Claude
Handles attached pages, so locating and reading can be tested together.
Last checked July 2026