Skip to content

Glossary

Success criteria

The specific results a trial has to produce to be judged worth continuing, settled before it begins rather than argued over afterwards.

In plain terms

What has to be true at the end for this to have been worth doing. Written down before you start, because afterwards everybody remembers wanting something slightly different, and whoever is most invested remembers most confidently.

01

Why it matters

Because a trial without them cannot fail. Something interesting always happened, somebody always liked it, and in the absence of an agreed test the decision falls to whoever argues best, which is reliably the person who wanted it in the first place.

02

How it works

They are about a decision rather than a deliverable. The question is whether to continue, expand or stop, which is different from whether a piece of work was finished properly, and confusing the two produces criteria that measure the wrong thing carefully.

The part that does the work is the stopping rule. Naming in advance the result that would make you stop is uncomfortable, is the reason criteria get written vaguely, and is the only thing that makes the rest of them binding.

A threshold has to be a number or a plain yes and no. If a result cannot be checked against the criterion by somebody who was not in the room, the criterion is a hope written in the grammar of a target.

They need a comparison to mean anything. A criterion stating a level with nothing to compare it against cannot distinguish the tool from everything else that changed, so the honest version names what the current position is and how it was measured.

Two or three is the working number. A list long enough to be thorough is long enough that some of it will be met and some will not, which returns the decision to interpretation and defeats the purpose of having written anything.

They have to survive the trial being interesting. A pilot always produces something nobody anticipated, and the temptation is to treat that discovery as the result, which quietly replaces the agreed test with a better story. Note the discovery, and judge against what was written.

Who agrees them decides whether they hold. Criteria written by the team running the trial and never shown to whoever pays for it will be revised at the end in perfect good faith, because nobody outside the room ever committed to them.

Criteria that decide, and criteria that decorate

Criteria that decide, and criteria that decorateNobody sets out to write the right-hand column. It appears because writing the left-hand one requires three uncomfortable admissions in a row: that you do not currently know what the number is, that you are prepared to name a level you might not reach, and that there is a result which would mean abandoning something people are enthusiastic about. Each of those is easier to postpone than to make, and the postponement is invisible because the resulting document looks like a plan. The tell is simple enough to apply to any trial plan in a minute. Read each criterion and ask what evidence would show it had not been met. If nothing would, the criterion is not doing any work, and a plan made entirely of such criteria will conclude, whatever happens, that the trial was a success and the next phase is warranted. That conclusion was written before the trial began; the trial merely supplied the anecdotes.DecidesA number, and where it istoday.A date by which it must hold.The result that would stopthis.DecoratesImprove efficiency.Positive feedback from theteam.Demonstrate the potential.The right-hand column appears inmost trial plans and cannotproduce a decision, becausenothing in it can turn outfalse. The left-hand columntakes an afternoon to write andhalf of that afternoon is thethird line.
Nobody sets out to write the right-hand column. It appears because writing the left-hand one requires three uncomfortable admissions in a row: that you do not currently know what the number is, that you are prepared to name a level you might not reach, and that there is a result which would mean abandoning something people are enthusiastic about. Each of those is easier to postpone than to make, and the postponement is invisible because the resulting document looks like a plan. The tell is simple enough to apply to any trial plan in a minute. Read each criterion and ask what evidence would show it had not been met. If nothing would, the criterion is not doing any work, and a plan made entirely of such criteria will conclude, whatever happens, that the trial was a success and the next phase is warranted. That conclusion was written before the trial began; the trial merely supplied the anecdotes.
03

Seen in the wild

  • A trial whose criterion was that people liked using it, which nobody could fail.

    ChatGPT
  • A search deployment judged against how long the same questions took before.

    Glean
  • Criteria written into the plan and visible to everyone from the start.

    ClickUp (Brain)
04

Common misconceptions

People assume

Everybody knows what success looks like.

In fact

Everybody has a version, and the versions differ in ways that only surface at the end. Writing them down converts a disagreement that would have happened later into one that happens now, when it is cheap.

People assume

Criteria are the thresholds.

In fact

The thresholds are the easy half. The part that makes them binding is the stopping rule, and criteria without one describe an aspiration rather than a decision anybody has agreed to make.

05

Questions

Why does a stopping rule matter so much?
Because without one the criteria only ever justify continuing. Naming the result that would end the work is what converts a list of hopes into a decision somebody has agreed in advance to accept, and it is the part people quietly leave out.
How many should we have?
Two or three. A longer list guarantees a mixed result, some met and some not, which hands the decision back to whoever argues most persuasively about which ones mattered, and that is precisely the situation writing anything down was meant to prevent.
How are these different from acceptance criteria?
These decide whether to carry on; acceptance criteria decide whether a piece of work is finished. A deliverable can meet its acceptance criteria completely and still fail the case for continuing, which is a perfectly coherent and quite common outcome.
06

Key takeaways

  • They govern a decision to continue, not whether work was finished.
  • The stopping rule is what makes the rest of them binding.
  • Two or three; a long list guarantees a result open to interpretation.
  • Whoever pays has to agree them, or they get revised at the end.
08

Tools that use this

  • ChatGPT

    A trial whose only criterion was that people liked it.

  • Glean

    Judged against how long the same questions took before.

  • ClickUp (Brain)

    Criteria in the plan and visible from the start.

Last checked August 2026

All glossary terms