Skip to content

Glossary

Cost per task

Cost per task is what one finished piece of work costs to produce, which is the only unit that lets AI spending be compared with the alternative ways of getting it done.

In plain terms

Not what the subscription costs and not what a month costs, but what it costs to get one thing done: one summarised document, one drafted reply, one processed record. That unit is the one that can be set against how else the work could happen, and it is the only figure in this area that is directly meaningful to somebody outside the discussion.

01

Why it matters

Because every other pricing unit is in the vendor's terms and this one is in yours. A per-person price, a bundle of credits and a token rate all describe what you are buying; cost per task describes what you are getting. It is also the only unit that travels: a colleague who does not know what a token is understands immediately what it costs to process one application, and that is the conversation where decisions get made.

02

How it works

Define the task as something finished, which is harder than it sounds and is where most attempts go wrong. One completed reply that could actually be sent, not one model response; one record processed and correct, not one call made. A definition that stops short of the finished thing produces a figure that looks good and cannot be compared with anything.

The first attempt is rarely the whole cost, and this is what separates a real figure from a flattering one. Retries after a poor answer, a second pass to check, the occasional case that falls back to a person: all of those belong in the number, spread across the tasks that needed them. A per-attempt cost is a different measurement wearing this one's name.

Human time still belongs in it wherever the output is checked. If somebody reads every draft before it goes out, that reading is part of what a completed task costs, and leaving it out is what produces figures that seem to make the labour comparison trivially favourable. Including it is also what makes the number credible to whoever is being asked to act on it.

It moves with volume, and usually downwards, which is worth stating when a first measurement looks discouraging. Early tasks carry the setup, the prompt refinement and the learning; the hundredth carries almost none of that. A figure from the first week is a poor guide to the steady state, and quoting it as though it were stable understates a good case.

The unit does not fit all work equally, and forcing it produces false precision. Repetitive bounded work has a natural task and a meaningful figure. Open-ended thinking, exploration and drafting where the value is in the quality rather than the throughput do not, and for those the honest position is that this unit does not apply rather than a number computed from an arbitrary denominator.

Two figures, both called cost per task

Two figures, both called cost per taskThe reason the left-hand figure is so persistent is that it is the easy one to obtain: it falls out of a billing page without anybody having to define what a completed task is. The right-hand figure requires deciding what finished means, tracking a batch through to that point, and being honest about the human minutes spent checking. That work is the entire value of the exercise, because it is what makes the number safe to set beside an alternative. It also tends to produce a result that is still favourable but considerably less dramatic, which is a feature: a figure that survives a sceptical question is worth more in the meeting than one that has to be defended and cannot be.The flattering oneModel spend divided byattempts.Retries not counted.Checking time not counted.The comparable oneBatch spend divided by taskscompleted.Retries and fallbacks included.Review time included.The left figure is genuinelyuseful for a different purpose,which is understanding what themodel costs. It is notcomparable with any alternativeway of doing the work, and it isusually the one quoted when acomparison is being made.
The reason the left-hand figure is so persistent is that it is the easy one to obtain: it falls out of a billing page without anybody having to define what a completed task is. The right-hand figure requires deciding what finished means, tracking a batch through to that point, and being honest about the human minutes spent checking. That work is the entire value of the exercise, because it is what makes the number safe to set beside an alternative. It also tends to produce a result that is still favourable but considerably less dramatic, which is a feature: a figure that survives a sceptical question is worth more in the meeting than one that has to be defended and cannot be.
03

Seen in the wild

  • What it costs to process one incoming record end to end, including the ones that need a second pass.

    n8n
  • What one summarised document costs, once the occasional re-run on a poor first answer is included.

    ChatGPT
  • What one answered internal question costs, against the alternative of somebody spending twenty minutes looking.

    Glean
04

Common misconceptions

People assume

It is the model cost divided by the number of tasks.

In fact

That is the cost per attempt. A real figure includes the retries, the checking pass and the cases that fell back to a person, spread across the tasks that completed. The two can differ substantially, and the gap is largest on exactly the work where the tool struggles most.

People assume

A low figure means the work should be automated.

In fact

It means the cost is low, which is one input. Whether the output is good enough unsupervised, what a mistake costs, and whether anybody would notice an error are separate questions, and a cheap task done wrong at volume is worse than an expensive one done right.

05

Telling them apart

Cost per task vs Inference cost

Cost per task

What one finished piece of work costs, including retries and checking.

Inference cost

What one call to the model costs.

If the number does not include the attempts that failed, it is the other one.

06

Questions

How do we calculate it?
Take a batch of real tasks through to completion, count everything spent across the batch including retries and any human checking, and divide by the number that actually finished. Measuring a batch rather than a single task is what captures the failures, and the failures are what distinguish this figure from a per-attempt one.
Should human review time be included?
Wherever it happens, yes. If every output is read before use, that reading is part of what a completed task costs, and omitting it produces a figure that makes the comparison look easy while being wrong. Including it is also what makes the number survive contact with anybody sceptical.
Why did our figure improve so much?
Early tasks carry the setup and the prompt refinement that later ones do not, so a measurement from the first week is not the steady state. That works in your favour, which is why it is worth re-measuring once the work has settled rather than reporting an early figure as though it were representative.
Does it work for every kind of work?
No, and forcing it produces false precision. Repetitive bounded work has a natural unit; open-ended drafting and exploration do not, and the value there is in quality rather than throughput. Saying the unit does not apply is more useful than a number computed against an arbitrary denominator.
07

Key takeaways

  • Define the task as something finished, not as one model response.
  • Include retries, checking and the cases that fell back to a person.
  • Measure a batch, because the failures are what make the figure real.
  • It falls with volume, so an early measurement understates a good case.
  • For open-ended work the honest answer is that the unit does not apply.
09

Tools that use this

  • n8n

    One record processed end to end, second passes included.

  • ChatGPT

    One summarised document, with the occasional re-run counted.

  • Glean

    One answered question, against twenty minutes of somebody looking.

Last checked July 2026

All glossary terms