Skip to content

Glossary

Acceptance criteria

The conditions a delivered piece of work must meet to count as finished, written so both sides can tell whether they hold.

In plain terms

What has to be true before you agree the work is finished. The useful test is whether somebody who was not in any of the meetings could look at the delivered thing and say yes or no without having to ask either side what they meant.

01

Why it matters

Because this is where money and goodwill are lost at the end of engagements that went well until then. Both sides believed they agreed, both were sincere, and the criterion turns out to admit two readings that only diverge once somebody wants to be paid.

02

How it works

The test is whether a stranger could adjudicate. If settling a disagreement requires asking either party what they had in mind, the criterion has not been written yet, however carefully both sides discussed it.

Adjectives are where the ambiguity hides. Accurate, robust, user-friendly and fast all feel specific in a room where everybody already agrees, and each of them is a placeholder for a measurement nobody has yet decided how to take.

They describe the deliverable rather than the decision to continue. Work can meet every criterion and still be something you would not commission again, and separating those two judgements keeps each of them honest.

Work whose output varies between runs breaks the usual form. A criterion naming what the system produces assumes it produces the same thing twice, which is not true of these tools, so it has to describe a standard held across a sample instead of a single result.

That means agreeing the sample as well as the standard. Which cases, chosen by whom, and judged by which rule: without those the criterion is unenforceable, because each side will demonstrate on the examples that suit it and both demonstrations will be honest.

Who judges is part of the criterion, not an administrative detail. A rule applied by two people produces two answers on the borderline cases, so naming one reviewer, or a way of settling a split, removes the last place a clear criterion can still fail.

The hard cases belong in the sample deliberately. A sample drawn from ordinary work will be passed by a system that fails on everything unusual, and the unusual cases are the ones the disagreement will eventually be about.

Two criteria about the same deliverable

Two criteria about the same deliverableThe right-hand column looks pedantic and is faster in practice, because it front-loads a disagreement that is otherwise deferred to the least convenient possible moment. Writing it takes an hour, and most of that hour is spent choosing the sample, which is where the real negotiation lives: a supplier will propose typical documents and a buyer should propose the awkward ones, and where the two land is a more informative conversation about the work than any amount of discussion about quality. The second right-hand line is worth copying directly, because it is the criterion that catches the failure specific to these tools. A summary that omits something is visibly incomplete and gets noticed; a summary that adds something plausible and absent reads perfectly and is the reason the work needed checking at all. Naming that as the standard, on an agreed sample, with one person deciding, converts an argument about quality into a morning's reading.Cannot be adjudicatedSummaries must be accurate.Handles our documents reliably.Output is of professionalquality.Can beOn these 40 documents, bothsides picked.No summary states somethingabsent.Judged by one named reviewer.Every line on the left wasagreed enthusiastically by bothparties at the time. Every lineon the right survives the momentsomebody wants to be paid, whichis the only moment the criteriaare ever read carefully.
The right-hand column looks pedantic and is faster in practice, because it front-loads a disagreement that is otherwise deferred to the least convenient possible moment. Writing it takes an hour, and most of that hour is spent choosing the sample, which is where the real negotiation lives: a supplier will propose typical documents and a buyer should propose the awkward ones, and where the two land is a more informative conversation about the work than any amount of discussion about quality. The second right-hand line is worth copying directly, because it is the criterion that catches the failure specific to these tools. A summary that omits something is visibly incomplete and gets noticed; a summary that adds something plausible and absent reads perfectly and is the reason the work needed checking at all. Naming that as the standard, on an agreed sample, with one person deciding, converts an argument about quality into a morning's reading.
03

Seen in the wild

  • A criterion saying summaries must be accurate, with nobody having defined accurate.

    ChatGPT
  • An agreed sample of awkward cases, judged by a rule both sides wrote down.

    Relevance AI
  • Criteria recorded against the work item rather than in somebody's inbox.

    ClickUp (Brain)
04

Common misconceptions

People assume

Both sides agreeing means the criteria are clear.

In fact

Agreement is easy while nothing is at stake, and the readings only diverge when somebody wants to be paid. The test is not whether you agree now but whether an outsider could settle it later without consulting either of you.

People assume

The same criteria work for AI deliverables.

In fact

Criteria written for software assume the same input produces the same output. Where it does not, a criterion about the result has to become a criterion about a standard held across an agreed sample, or it cannot be checked at all.

05

Questions

What is the quickest way to test a criterion?
Read it and ask whether somebody who has never met either party could decide, from the delivered work alone, whether it holds. If settling it would need a conversation about what was meant, the criterion is still a discussion rather than a criterion.
How do we write these for output that varies?
Describe a standard held over an agreed sample rather than a single result. That means naming the cases, who picks them, and the rule for judging each one, because otherwise both sides will demonstrate on the examples that suit them.
How do these differ from success criteria?
These say whether a deliverable is finished; success criteria say whether the wider effort is worth continuing. Work can satisfy every acceptance criterion and still not justify a second phase, and keeping the two judgements apart is what stops either of them being quietly bent to suit the other.
06

Key takeaways

  • The test is whether a stranger could adjudicate without asking you.
  • Adjectives are placeholders for measurements nobody has agreed.
  • Varying output needs a standard over a sample, not a single result.
  • Put the awkward cases in the sample; that is what disputes are about.
08

Tools that use this

  • ChatGPT

    A criterion requiring accuracy that nobody has defined.

  • Relevance AI

    An agreed sample of awkward cases with a written rule.

  • ClickUp (Brain)

    Criteria kept against the work item, not in an inbox.

Last checked August 2026

All glossary terms