Skip to content

Glossary

Token pricing

Token pricing charges for the amount of text going in and coming out rather than for the person doing the asking, which makes the length of the work the thing that costs money.

In plain terms

Everything you send is chopped into small pieces and everything that comes back is too, and you pay by the piece. A short question about a paragraph costs very little. The same question about a hundred-page contract costs considerably more, because the contract went in as well. Once that is clear, most of the surprising things about AI bills stop being surprising.

01

Why it matters

Because it is the only pricing model in common use where you can genuinely compare two vendors, and because it is the layer everything else is built on. Credits bundle it, seats hide it, and allowances ration it, but underneath the arithmetic is the same. A buyer who understands this layer can work out what any of the others is really offering, and a buyer who does not is comparing headline numbers in units nobody has defined.

02

How it works

Most buyers never meet this pricing directly, which is worth establishing before anything else. Where a product charges a subscription or a bundle, the vendor is paying by the token and selling you something else on top, so this page describes their costs rather than your invoice. Where you supply your own provider key, or use a platform that passes the charge through, you are on this meter yourself. The mechanism is worth understanding either way, because it explains what a bundle is priced against.

Input and output are charged at different rates, and output is usually the dearer of the two. That single fact explains a large share of unexpected bills: a task that reads a great deal and answers briefly behaves quite differently from one that reads a little and writes at length, even though both look like one request. The reason is that reading can be processed in bulk while writing has to proceed one piece at a time, each one depending on what came before, so producing text occupies the hardware for longer than receiving it does.

Everything sent counts, including the parts nobody typed. The instructions the product adds behind the scenes, the documents attached for grounding, the earlier turns of a conversation being resent so the model can follow it: all of that is input and all of it is charged. This is the commonest reason a bill exceeds what the visible work seems to justify.

A long conversation gets more expensive as it continues, which surprises people because nothing about it feels different. Each new question carries the whole exchange so far, so the input grows with every turn. Starting a fresh conversation for an unrelated question is the cheapest habit available in this whole subject.

Different models charge very differently for the same work, sometimes by a wide margin, and the more capable one is not always the right choice. Where a task is routine, a cheaper model doing it adequately is a real saving repeated at volume; where it is hard, the cheaper model's rework can cost more than the difference. Deciding this per task rather than once for everything is where most of the available saving actually sits.

Reasoning output is charged even where it is not shown to you. Models that work through a problem before answering produce that working as output, and the fact that a product may hide it does not make it free. A task that looks like a two-line answer can carry a great deal of invisible generation behind it.

The rate is not the bill, and the difference is volume. A figure that looks trivially small per request is being multiplied by however many requests an automated process makes, which is where a comfortable experiment becomes an uncomfortable invoice. Estimating from the shape of your work rather than from the rate is the only reliable approach.

Work that is not urgent can often be charged less, which is the lever most commonly left unused. Providers frequently offer a reduced rate for requests submitted to be processed when there is capacity rather than immediately, and a great deal of automated work has no one waiting on it: overnight processing, bulk classification, anything whose output is read the following morning. Distinguishing the work that genuinely needs an answer now from the work that merely gets one is where this saving is found.

Rates move over time and the direction is not uniform, which affects planning more than most buyers expect. The long trend has been downward, and individual prices still rise: an introductory rate ends, or a model is repriced upward at a scheduled date. A decision made on today's relative costs between two models can be wrong within a year in either direction, so an architecture that makes swapping the model easy is worth something specific, and a long commitment made on current rates is worth examining carefully.

What the user sees, what is charged

What the user sees, what is chargedThis gap is worth dwelling on because it is the single most useful thing to understand about paying for AI, and because nothing in the interface hints at it. The product is designed to make the exchange feel like a conversation, and a conversation that displayed its own system instructions, re-sent history and internal working would be unusable. So the design is right and the consequence is a genuine mismatch between what a user experiences and what an invoice reflects. The practical upshot is not to distrust the bill but to stop estimating from the visible side. Anybody forecasting spend should be reasoning about the material going in, including the parts nobody typed, and about how many times a day that happens, which is a different exercise from thinking about how many questions people ask.The conversationOne short question.One short answer.Feels like a small amount ofwork.What passed throughHidden instructions, plus anattached document.The whole earlier exchange,resent.Reasoning generated but notdisplayed.Both descriptions are of thesame request. The gap betweenthem is where nearly everysurprising AI bill comes from,and none of it is concealment:it is simply not shown, becauseshowing it would be unhelpful tothe person having theconversation.
This gap is worth dwelling on because it is the single most useful thing to understand about paying for AI, and because nothing in the interface hints at it. The product is designed to make the exchange feel like a conversation, and a conversation that displayed its own system instructions, re-sent history and internal working would be unusable. So the design is right and the consequence is a genuine mismatch between what a user experiences and what an invoice reflects. The practical upshot is not to distrust the bill but to stop estimating from the visible side. Anybody forecasting spend should be reasoning about the material going in, including the parts nobody typed, and about how many times a day that happens, which is a different exercise from thinking about how many questions people ask.
03

Seen in the wild

  • Attaching a long document to a question, where the document is charged as input before any answer is produced.

    ChatGPT
  • Comparing rates across several providers in one place, which is the clearest way to see how widely they differ for the same task.

    OpenRouter
  • An automation calling a model on every incoming record, where a small per-request figure is multiplied by the volume.

    n8n
  • Running a model on your own hardware, where the cost becomes electricity and equipment rather than a per-token rate.

    Ollama
04

Common misconceptions

People assume

A token is a word.

In fact

It is a fragment, and the ratio varies by language and by content: ordinary English runs somewhat more tokens than words, while code, unusual names and other languages can run considerably higher. For estimating purposes the useful habit is to think in rough proportion to length rather than to count anything precisely.

People assume

The cheap model is the cheap option.

In fact

Only where it does the job. A cheaper model that needs its work checked or redone can cost more in total than the capable one that got it right, and the effect is largest on exactly the difficult tasks where the saving looks most attractive. The decision belongs to the task rather than to the organisation.

People assume

We are not charged for what we cannot see.

In fact

System instructions, attached documents, resent conversation history and hidden reasoning are all charged, and together they frequently exceed what the user typed. The visible exchange is a poor guide to the bill, which is why bills feel disproportionate to people who have only seen the conversation.

05

Telling them apart

Token pricing vs Seat pricing

Token pricing

Charges for work done. Comparable between vendors, harder to forecast.

Seat pricing

Charges for people. Easy to forecast, indifferent to how much they do.

If a quiet month costs less, you are on tokens.

06

Questions

Why is our bill higher than the conversation suggests?
Because most of what is charged was never typed. Behind-the-scenes instructions, attached documents, the earlier turns being resent each time, and any hidden reasoning all count as input or output. The visible exchange is usually a small part of what passed through, which is why the two feel unrelated.
Why does a long conversation get more expensive?
Each question carries the whole exchange so far, so the input grows with every turn while the questions stay the same size. Nothing feels different, which is why it goes unnoticed. Starting a new conversation when the subject changes is the simplest saving available anywhere in this subject.
Should we use a cheaper model?
For routine work, usually, and the saving is real at volume. For difficult work it often costs more once rework is counted, because the tasks where a cheap model struggles are the ones where checking and redoing are expensive. The useful version of this decision is made per task, not once for everything.
How do we estimate before committing?
From the shape of your work rather than from the rate. Take a representative task, note roughly how much goes in and how much comes back, and multiply by how often it happens. That is a cruder method than it sounds and it is far more reliable than reasoning from a per-unit figure, which gives no sense of volume.
Why does this page quote no prices?
Because rates change frequently and an out-of-date figure in a reference page is worse than none: it will be believed. The mechanism is stable and the numbers are not, so what is written here is the arithmetic, and the current rates belong on the vendor's own pricing page where they are maintained.
Does self-hosting avoid this?
It replaces the meter with fixed costs rather than removing cost. Hardware, electricity and the attention of whoever keeps it running are the price instead, which suits steady predictable volume and suits occasional use badly. The comparison worth making is against your actual pattern, not against the per-token rate.
07

Key takeaways

  • Input and output are charged differently, and output is usually dearer.
  • Most of what you pay for was never typed: instructions, attachments, resent history, hidden reasoning.
  • A long conversation grows more expensive with every turn.
  • Model choice belongs to the task; the cheap model is only cheap where it works.
  • Estimate from the shape of your work, never from the per-unit rate.
09

Tools that use this

  • ChatGPT

    A long attachment charged as input before any answer appears.

  • OpenRouter

    Rates across several providers in one place, where the spread is visible.

  • Ollama

    The alternative shape: equipment and electricity instead of a meter.

Last checked July 2026

All glossary terms