Skip to content

Glossary

Vision language model (VLM)

A vision language model takes pictures and words in the same request, so a screenshot, a chart or a photograph can be asked about in ordinary language rather than described first.

In plain terms

It looks at a picture and talks about it. Send a photograph of a form, a screenshot of an error, a chart from a report, and ask a question in ordinary words. What matters is not that it can see but what it does when the picture is bad, because a blurry corner produces an answer just as confident as a clear one.

01

Why it matters

Because it removes the step where somebody had to describe the picture first. A screenshot of an error, a photograph of an invoice or a chart nobody has the underlying figures for can all now be asked about directly, which converts a great deal of tedious transcription into a question. The risk arrives in the same motion.

02

How it works

Pictures and words are handled in one request rather than by two systems in sequence. There is no separate step that converts an image to text and hands it on, so a question can refer to what is in the picture and what is in the sentence at the same time, which is why asking about a specific part of a screenshot works at all.

It reads a picture as a whole rather than transcribing it, and that difference cuts both ways. Layout, emphasis, what a chart is doing and where a form's fields sit are all available, while an exact string of small print is less reliable than a dedicated character reader would be. Meaning is the strength; literal accuracy is not.

A poor image produces a confident answer, which is the failure worth planning for. Glare, a cropped edge, a low-resolution scan or a handwritten note all degrade what can be read, and nothing about the response signals that degradation, so an invented figure arrives in the same tone as a correct one. The picture quality is the control.

Its everyday value is unglamorous and large. Asking about an error on screen, pulling the numbers out of a chart in a slide, checking a photograph of a document against a rule: none of these is impressive to demonstrate and all of them replace a person retyping something. That is where the hours actually are.

Two questions a picture can answer

Two questions a picture can answerThe reason people end up disappointed here is that a single photograph usually contains both columns and the request does not distinguish them. Somebody photographs an invoice and asks what it says, meaning partly what is this document and partly what is the exact total, and receives an answer that is excellent on the first and quietly approximate on the second. Nothing in the response separates the two. The habit that fixes it costs nothing: decide before asking whether a wrong character would matter. If it would, the number needs checking against the image by a person or by a dedicated reader, and the general model should be asked about everything except that. If it would not, this is the better tool by a wide margin and the exact reader would have produced a wall of text nobody wanted.What does it say?Exact characters.A reference, a code, a total.A wrong letter is a defect.A dedicated reader wins.What is going on?Layout and emphasis.What a chart is showing.Which field is which.This is where it wins.Both look like reading a pictureand they are graded on oppositethings. Choosing by which columnyour task sits in settles thetool question immediately.
The reason people end up disappointed here is that a single photograph usually contains both columns and the request does not distinguish them. Somebody photographs an invoice and asks what it says, meaning partly what is this document and partly what is the exact total, and receives an answer that is excellent on the first and quietly approximate on the second. Nothing in the response separates the two. The habit that fixes it costs nothing: decide before asking whether a wrong character would matter. If it would, the number needs checking against the image by a person or by a dedicated reader, and the general model should be asked about everything except that. If it would not, this is the better tool by a wide margin and the exact reader would have produced a wall of text nobody wanted.
03

Seen in the wild

  • A general assistant that works with text, documents, spreadsheets and images in the same conversation rather than requiring each in its own tool.

    ChatGPT
  • An assistant answering with images and files understood in one session, holding very long documents alongside them.

    Google Gemini
  • A tool that watches a process being performed and produces a written guide with annotated screenshots, which is reading a screen for meaning rather than for text.

    Scribe
04

Common misconceptions

People assume

It is character recognition with a chat window.

In fact

It reads for meaning rather than for exact characters, which makes it better at layout, charts and what a document is doing and worse at reproducing small print verbatim. Where an exact string matters, a dedicated reader remains the right tool and this is the wrong one.

People assume

A bad photograph will produce an obviously bad answer.

In fact

It produces a confident one. Glare, cropping and low resolution all reduce what can actually be read while changing nothing about the tone of the response, so the degradation is invisible in the output. Controlling the image is the only place that risk can be managed.

05

Telling them apart

Vision language model vs OCR

Vision language model

Understands what the picture is doing.

OCR

Converts the characters into text, exactly.

Ask which you need by asking whether a wrong character matters. For a contract reference it does; for a question about a chart it does not.

06

Questions

Can it replace character recognition?
Not where exactness matters, and it beats it comfortably where meaning does. A reference number, a legal citation or a figure copied into an invoice all need character-level accuracy that a dedicated reader provides. Understanding what a chart shows or where a form's sections sit is the opposite case, and the wrong tool there is the exact reader.
How do I stop it inventing details from a poor image?
By controlling the image rather than the request. Better light, a straight angle, a full frame and a higher resolution all raise what can genuinely be read, and none of the phrasing tricks do. Asking it to say when something is illegible helps a little and is far less reliable than simply supplying a picture it can read.
Where does it pay for itself?
In the transcription nobody enjoys. Reading numbers off a chart that arrived as an image, checking a photographed document against a rule, or explaining an error from a screenshot all replace somebody retyping or describing. None of it demonstrates well and it accounts for most of the practical value.
07

Key takeaways

  • Pictures and words arrive in one request, so no describing step is needed.
  • It reads for meaning, which beats exact transcription and loses to it on small print.
  • A poor image degrades the answer without changing its confidence.
  • The value is in ordinary transcription work, not in the demonstrations.
09

Tools that use this

  • ChatGPT

    Text, documents, spreadsheets and images in one conversation.

  • Google Gemini

    Images and files understood alongside very long documents.

  • Scribe

    Reads a screen for meaning and writes the guide from it.

Last checked July 2026

All glossary terms