Glossary
Acceptance criteria
The conditions a delivered piece of work must meet to count as finished, written so both sides can tell whether they hold.
In plain terms
What has to be true before you agree the work is finished. The useful test is whether somebody who was not in any of the meetings could look at the delivered thing and say yes or no without having to ask either side what they meant.
Why it matters
Because this is where money and goodwill are lost at the end of engagements that went well until then. Both sides believed they agreed, both were sincere, and the criterion turns out to admit two readings that only diverge once somebody wants to be paid.
How it works
The test is whether a stranger could adjudicate. If settling a disagreement requires asking either party what they had in mind, the criterion has not been written yet, however carefully both sides discussed it.
Adjectives are where the ambiguity hides. Accurate, robust, user-friendly and fast all feel specific in a room where everybody already agrees, and each of them is a placeholder for a measurement nobody has yet decided how to take.
They describe the deliverable rather than the decision to continue. Work can meet every criterion and still be something you would not commission again, and separating those two judgements keeps each of them honest.
Work whose output varies between runs breaks the usual form. A criterion naming what the system produces assumes it produces the same thing twice, which is not true of these tools, so it has to describe a standard held across a sample instead of a single result.
That means agreeing the sample as well as the standard. Which cases, chosen by whom, and judged by which rule: without those the criterion is unenforceable, because each side will demonstrate on the examples that suit it and both demonstrations will be honest.
Who judges is part of the criterion, not an administrative detail. A rule applied by two people produces two answers on the borderline cases, so naming one reviewer, or a way of settling a split, removes the last place a clear criterion can still fail.
The hard cases belong in the sample deliberately. A sample drawn from ordinary work will be passed by a system that fails on everything unusual, and the unusual cases are the ones the disagreement will eventually be about.
Two criteria about the same deliverable
Seen in the wild
A criterion saying summaries must be accurate, with nobody having defined accurate.
ChatGPTAn agreed sample of awkward cases, judged by a rule both sides wrote down.
Relevance AICriteria recorded against the work item rather than in somebody's inbox.
ClickUp (Brain)
Common misconceptions
People assume
Both sides agreeing means the criteria are clear.
In fact
Agreement is easy while nothing is at stake, and the readings only diverge when somebody wants to be paid. The test is not whether you agree now but whether an outsider could settle it later without consulting either of you.
People assume
The same criteria work for AI deliverables.
In fact
Criteria written for software assume the same input produces the same output. Where it does not, a criterion about the result has to become a criterion about a standard held across an agreed sample, or it cannot be checked at all.
Questions
- What is the quickest way to test a criterion?
- Read it and ask whether somebody who has never met either party could decide, from the delivered work alone, whether it holds. If settling it would need a conversation about what was meant, the criterion is still a discussion rather than a criterion.
- How do we write these for output that varies?
- Describe a standard held over an agreed sample rather than a single result. That means naming the cases, who picks them, and the rule for judging each one, because otherwise both sides will demonstrate on the examples that suit them.
- How do these differ from success criteria?
- These say whether a deliverable is finished; success criteria say whether the wider effort is worth continuing. Work can satisfy every acceptance criterion and still not justify a second phase, and keeping the two judgements apart is what stops either of them being quietly bent to suit the other.
Key takeaways
- The test is whether a stranger could adjudicate without asking you.
- Adjectives are placeholders for measurements nobody has agreed.
- Varying output needs a standard over a sample, not a single result.
- Put the awkward cases in the sample; that is what disputes are about.
Tools that use this
- ChatGPT
A criterion requiring accuracy that nobody has defined.
- Relevance AI
An agreed sample of awkward cases with a written rule.
- ClickUp (Brain)
Criteria kept against the work item, not in an inbox.
Last checked August 2026