Skip to content

Glossary

Context pricing

Everything sent to a model is charged as it goes in, so the size of what you attach drives the cost of asking about it more than the question does.

In plain terms

Ask a short question about a hundred-page report and the hundred pages are what you are mostly paying for, not the question. The model has to receive the material to answer about it, and receiving is charged. It follows that the same short question costs wildly different amounts depending on what was attached to it.

01

Why it matters

Because it inverts the intuition people bring from every other tool, where the effort tracks what you asked rather than what you supplied. Once the cost is understood as attaching to the material, the pattern of surprising bills becomes predictable and, more usefully, addressable: the lever is what gets sent, which is under your control in a way the vendor's rate is not.

02

How it works

The charge repeats on every question rather than being paid once for the document. Asking six things about the same report generally sends the report six times, because each request stands alone unless the product has arranged otherwise. That is the single most surprising property here, and it explains bills that seem out of proportion to a short session.

Larger capacity is an option rather than an obligation, which is worth saying because the marketing runs the other way. A model that accepts very large inputs makes some work possible and it charges for what you actually send, so the capacity itself costs nothing until used. The trap is treating a large window as an invitation to attach everything.

Sending the relevant part is usually both cheaper and better. Retrieving the three pages that bear on a question rather than attaching the whole file reduces what is charged and tends to improve the answer, because there is less irrelevant material to work through. It is one of the few places where the economical choice and the good one point the same way.

Where the same material is sent repeatedly, caching is the mechanism that addresses this specifically, and it is worth knowing before optimising anything else. It is not universal and its terms vary, but where a long fixed preamble accompanies every request it changes the arithmetic more than anything a user can do by hand.

Six questions about one report

Six questions about one reportThis is the clearest example of the general pattern that what is charged and what is experienced come apart. From the user's side there was one document and a conversation about it, which is exactly the interaction the interface was designed to provide. Underneath, each question was an independent request that had to carry the whole report in order to be answerable at all. Understanding that changes behaviour in a useful direction rather than a fearful one: it argues for asking a document several things at once rather than in sequence, for attaching the section that matters rather than the file, and for finding out whether a product caches repeated material. All three are things a user can do, which makes this one of the more actionable ideas in the pricing category.What it feels likeOne upload.Six short questions.A small amount of work.What was sentThe report, six times.Six short questions.Nearly all of it the report.Nothing here is a fault in theproduct. A request carries whatthe model needs to answer it,and the model does not rememberthe previous one unlesssomething has been built to makeit appear to.
This is the clearest example of the general pattern that what is charged and what is experienced come apart. From the user's side there was one document and a conversation about it, which is exactly the interaction the interface was designed to provide. Underneath, each question was an independent request that had to carry the whole report in order to be answerable at all. Understanding that changes behaviour in a useful direction rather than a fearful one: it argues for asking a document several things at once rather than in sequence, for attaching the section that matters rather than the file, and for finding out whether a product caches repeated material. All three are things a user can do, which makes this one of the more actionable ideas in the pricing category.
03

Seen in the wild

  • Attaching a long report and asking several questions about it, where the report is charged on each one.

    ChatGPT
  • A grounded tool that retrieves the relevant passages rather than sending an entire corpus with every question.

    NotebookLM
  • Comparing models whose input capacity and input rates differ, where the two together decide what a long document costs.

    OpenRouter
04

Common misconceptions

People assume

We upload the document once, so we pay for it once.

In fact

Each question generally resends it, because a request carries everything the model needs to answer. Unless the product is caching or retrieving selectively, six questions about one report means the report going through six times, which is the usual explanation for a bill that does not match a short session.

People assume

A bigger context window costs more.

In fact

The capacity is not the charge; what you send is. A model that accepts very large inputs costs nothing extra while you send short ones. What the larger window changes is what becomes possible, and the risk it introduces is the habit of attaching everything because you now can.

05

Telling them apart

Context pricing vs Context window

Context pricing

What sending the material costs, each time you send it.

Context window

How much material the model can accept at once.

One is a limit and the other is a meter. Reaching the limit is an error; passing the meter is an invoice.

06

Questions

Why did a short session cost so much?
Almost always because something long was attached and resent with every question. The visible exchange was short; what passed through was the document multiplied by the number of questions. Asking fewer, better-combined questions about a large attachment is a genuine saving rather than a marginal one.
How do we reduce it?
Send less, by retrieving the relevant part rather than the whole file, and combine questions about one document into a single request where that is practical. Where the same long material accompanies every request, caching addresses it more effectively than anything done by hand.
Should we choose a model with a smaller window?
Not for cost reasons, since the capacity is not what is charged. Choose on whether the work needs the room, and manage the cost through what you actually send. A smaller window will make oversized attachments fail rather than making them cheap.
07

Key takeaways

  • The material is charged on the way in, so what you attach drives the cost.
  • It repeats on every question unless something is caching or retrieving.
  • Capacity is not a charge; a large window costs nothing until you fill it.
  • Sending the relevant part is usually cheaper and produces a better answer.
09

Tools that use this

  • ChatGPT

    A long report charged again on each question asked about it.

  • NotebookLM

    Retrieving the relevant passages instead of resending a corpus.

  • OpenRouter

    Where input capacity and input rate can be compared together.

Last checked July 2026

All glossary terms