Glossary
Cost per task
Cost per task is what one finished piece of work costs to produce, which is the only unit that lets AI spending be compared with the alternative ways of getting it done.
In plain terms
Not what the subscription costs and not what a month costs, but what it costs to get one thing done: one summarised document, one drafted reply, one processed record. That unit is the one that can be set against how else the work could happen, and it is the only figure in this area that is directly meaningful to somebody outside the discussion.
Why it matters
Because every other pricing unit is in the vendor's terms and this one is in yours. A per-person price, a bundle of credits and a token rate all describe what you are buying; cost per task describes what you are getting. It is also the only unit that travels: a colleague who does not know what a token is understands immediately what it costs to process one application, and that is the conversation where decisions get made.
How it works
Define the task as something finished, which is harder than it sounds and is where most attempts go wrong. One completed reply that could actually be sent, not one model response; one record processed and correct, not one call made. A definition that stops short of the finished thing produces a figure that looks good and cannot be compared with anything.
The first attempt is rarely the whole cost, and this is what separates a real figure from a flattering one. Retries after a poor answer, a second pass to check, the occasional case that falls back to a person: all of those belong in the number, spread across the tasks that needed them. A per-attempt cost is a different measurement wearing this one's name.
Human time still belongs in it wherever the output is checked. If somebody reads every draft before it goes out, that reading is part of what a completed task costs, and leaving it out is what produces figures that seem to make the labour comparison trivially favourable. Including it is also what makes the number credible to whoever is being asked to act on it.
It moves with volume, and usually downwards, which is worth stating when a first measurement looks discouraging. Early tasks carry the setup, the prompt refinement and the learning; the hundredth carries almost none of that. A figure from the first week is a poor guide to the steady state, and quoting it as though it were stable understates a good case.
The unit does not fit all work equally, and forcing it produces false precision. Repetitive bounded work has a natural task and a meaningful figure. Open-ended thinking, exploration and drafting where the value is in the quality rather than the throughput do not, and for those the honest position is that this unit does not apply rather than a number computed from an arbitrary denominator.
Two figures, both called cost per task
Seen in the wild
What it costs to process one incoming record end to end, including the ones that need a second pass.
n8nWhat one summarised document costs, once the occasional re-run on a poor first answer is included.
ChatGPTWhat one answered internal question costs, against the alternative of somebody spending twenty minutes looking.
Glean
Common misconceptions
People assume
It is the model cost divided by the number of tasks.
In fact
That is the cost per attempt. A real figure includes the retries, the checking pass and the cases that fell back to a person, spread across the tasks that completed. The two can differ substantially, and the gap is largest on exactly the work where the tool struggles most.
People assume
A low figure means the work should be automated.
In fact
It means the cost is low, which is one input. Whether the output is good enough unsupervised, what a mistake costs, and whether anybody would notice an error are separate questions, and a cheap task done wrong at volume is worse than an expensive one done right.
Telling them apart
Cost per task vs Inference cost
Cost per task
What one finished piece of work costs, including retries and checking.
What one call to the model costs.
If the number does not include the attempts that failed, it is the other one.
Questions
- How do we calculate it?
- Take a batch of real tasks through to completion, count everything spent across the batch including retries and any human checking, and divide by the number that actually finished. Measuring a batch rather than a single task is what captures the failures, and the failures are what distinguish this figure from a per-attempt one.
- Should human review time be included?
- Wherever it happens, yes. If every output is read before use, that reading is part of what a completed task costs, and omitting it produces a figure that makes the comparison look easy while being wrong. Including it is also what makes the number survive contact with anybody sceptical.
- Why did our figure improve so much?
- Early tasks carry the setup and the prompt refinement that later ones do not, so a measurement from the first week is not the steady state. That works in your favour, which is why it is worth re-measuring once the work has settled rather than reporting an early figure as though it were representative.
- Does it work for every kind of work?
- No, and forcing it produces false precision. Repetitive bounded work has a natural unit; open-ended drafting and exploration do not, and the value there is in quality rather than throughput. Saying the unit does not apply is more useful than a number computed against an arbitrary denominator.
Key takeaways
- Define the task as something finished, not as one model response.
- Include retries, checking and the cases that fell back to a person.
- Measure a batch, because the failures are what make the figure real.
- It falls with volume, so an early measurement understates a good case.
- For open-ended work the honest answer is that the unit does not apply.
Last checked July 2026