Glossary
Reasoning tokens
Reasoning tokens are the working a model generates on its way to an answer, billed and counted exactly like the reply itself even where the product never shows you a word of it.
In plain terms
You are billed for text you will never read. A model that works a problem through produces that working as ordinary output, and most products present a tidied summary or nothing at all while charging for the whole of it. This is the usual cause of an invoice that looks several times larger than the visible answers would suggest.
Why it matters
It is the specific thing that makes a first bill surprising. Somebody compares the length of the replies against the charge, cannot reconcile them, and concludes the pricing is wrong. The pricing is not wrong; the working is simply not on the screen. Once that is understood, the effort setting most products expose stops being an obscure preference and becomes the most direct control anybody has over the number.
How it works
The working is generated exactly as any other text is, one piece at a time, and priced as output. Whether it is displayed in full, summarised or hidden entirely is a decision about the interface. None of those choices changes what was produced or what it cost.
It also occupies the fixed budget the request has to fit inside, so a long stretch of working competes with your documents and your conversation for the same space. On a hard problem with a large attachment, that competition is real and shows up as an answer that seems to have lost track of something you supplied.
It accounts for most of the wait as well as most of the cost, which is why a considered reply takes noticeably longer while producing a visible answer no longer than a direct one. When a product feels slow on a hard question, the time is going here rather than into anything you can see.
Where a product exposes an effort or budget setting, that is a direct control over how much of this gets produced. Turning it down on routine work and up on the few genuinely hard problems is the ordinary way to run this, and it is more effective than any other cost adjustment available on a reasoning model.
One considered answer, as an illustration: what you read against what you paid for
Seen in the wild
Send the same problem twice through a routing interface, once to a deliberate model and once to a direct one, and compare the reported cost rather than the answers.
OpenRouterAsk a hard question of an assistant that shows its working and notice how much of it appears before the first line of the actual answer.
ClaudePut a considered model into an automation running at volume and check the first week's charge against what the visible outputs would suggest.
n8n
Common misconceptions
People assume
If we cannot see it, we are not paying for it.
In fact
Hiding it is a presentation choice and the charge is unaffected. Few misunderstandings in the cost vocabulary are more expensive, because it makes a bill look inexplicable and sends people investigating pricing rather than adjusting the setting that would fix it.
People assume
It only affects cost.
In fact
It occupies the same fixed space as your material, so a long stretch of working leaves less room for the documents and the conversation. That is why a hard question with a large attachment can produce an answer that has plainly lost track of something you supplied.
Telling them apart
Reasoning tokens vs Chain-of-thought
Reasoning tokens
The accounting view: what the working consumed, in budget, in time and in the space your material needed.
The same text seen as an artefact: what the steps say, and what they are useful for.
One question is what the working means and the other is what it cost. They are the same text, and only one of the two turns up on an invoice.
Questions
- Why is our bill so much larger than the answers look?
- Almost certainly this. The visible reply is a fraction of what was generated on a deliberate model, and the rest was working the product chose not to display. Compare the same task on a direct model for a day and the difference will be immediate.
- Can we reduce it without losing the benefit?
- Yes, by using the effort setting where a product offers one and reserving high settings for problems that genuinely have steps. Most requests in an ordinary workload have nothing to deliberate about, and sending those to a deliberate model at full effort is the usual reason a bill runs away.
- Does it count against usage limits as well?
- Generally yes, which is why a plan's allowance for a capable model runs out faster than the message count alone would explain. It is worth telling colleagues, because otherwise reaching the ceiling early reads as a fault rather than as the design working.
Key takeaways
- The working is ordinary output, charged and counted whether or not you ever see it.
- Hiding it is an interface decision that changes nothing about the cost.
- It occupies the same fixed space as your documents and conversation.
- It accounts for most of the extra wait as well as most of the extra cost.
- The effort setting is the most direct control over the number.
Tools that use this
- OpenRouter
Reported cost on the same problem, deliberate against direct.
- Claude
Working shown before the answer, so the volume of it is visible.
- n8n
At volume, where a first week's charge makes the point unmistakably.
Last checked July 2026