Skip to content

Glossary

Benchmark

A benchmark is a standard test scoring models against each other on a fixed set of questions, which makes it useful for drawing up a shortlist and unreliable as a prediction about your own work.

In plain terms

It is an exam with a published paper. Everybody sits the same questions and the marks get compared. That works well for spotting who is hopeless and works badly for predicting who will be good at your particular job, partly because your job is not on the paper and partly because by now everybody has seen the paper.

01

Why it matters

Scores drive purchasing decisions, press coverage and a great deal of internal argument, and the gap between scoring well and working for you is where the money goes. The useful reframing is that you are allowed to write your own exam. Twenty real cases from your own work, with answers somebody competent agrees are good, will tell you more in an afternoon than every published table put together, and it keeps telling you when the models change underneath.

02

How it works

A fixed set of questions with known answers is run identically across models and the results are scored, producing a number that can be put in a table. The value of the arrangement is comparability: whatever else is true, every model faced the same paper, which is more than can be said for most claims in this market.

Published tests leak into training material, and this is the central problem rather than an edge case. A model trained after a benchmark became public may have encountered its questions and answers, which turns a test of ability into a test of recall. Nobody can fully verify whether it happened, which is why a high score on a well-known test is weaker evidence than it looks.

Scores are frequently close and the ordering is less stable than the reporting suggests. Small differences fall within the variation you would get from running the same test twice, and a lead that changes hands between two releases was probably never a lead. Treat a table as a rough grouping into tiers rather than as a ranking.

What gets measured is narrow, because it has to be automatically checkable. Exam-style questions with definite answers work well. Tone, judgement, following a house style, knowing when to decline and being pleasant to work with over an hour are largely unmeasured, and for a great many business uses those are the whole of what you are buying.

Your own set is the answer and it is less work than it sounds. Twenty to fifty real cases from your actual work, with agreed good answers, run whenever a model or a prompt changes. It is the same discipline that makes prompt work repeatable, and it is the only score that has any bearing on whether the thing will help you.

Two exams, and which one is about you

Two exams, and which one is about youBoth columns produce a number and only one of them is about your organisation. A published benchmark is genuinely useful for the coarse question of which models are in contention at all, and it answers nothing finer than that: the questions came from somewhere else, the answers may have reached the model during training, and the differences between neighbouring entries are often smaller than the variation between two runs. Your own set inverts every one of those. The cases are ones you actually receive, the standard was set by the person who does the work, and nothing has leaked because nothing was published. The reason to build it is not purity. It is that a provider will change the model under your product at some point without asking, and a set you can rerun is how you find out what that did.A published benchmarkQuestions nobody at your firmwrote.Answers may be in the trainingmaterial.One number, often within noise.Marks what a machine can mark.Twenty of your own casesQuestions you actually receive.Answers agreed by whoever doesthe work.Rerun whenever anythingchanges.Marks what you are buying itfor.The left column is worth tenminutes for narrowing a longlist to three candidates. Theright column is what decides, itcosts an afternoon to build, andit keeps working every time aprovider changes somethingunderneath you.
Both columns produce a number and only one of them is about your organisation. A published benchmark is genuinely useful for the coarse question of which models are in contention at all, and it answers nothing finer than that: the questions came from somewhere else, the answers may have reached the model during training, and the differences between neighbouring entries are often smaller than the variation between two runs. Your own set inverts every one of those. The cases are ones you actually receive, the standard was set by the person who does the work, and nothing has leaked because nothing was published. The reason to build it is not purity. It is that a provider will change the model under your product at some point without asking, and a set you can rerun is how you find out what that did.
03

Seen in the wild

  • Run twenty of your own real cases across several models through one routing interface, which is the comparison that actually predicts something about your work.

    OpenRouter
  • Take the five questions you genuinely need answered well and put them to an assistant before reading a single published score.

    Claude
  • Ask the same question in two assistants and notice how often your own preference disagrees with the published ranking.

    ChatGPT
04

Common misconceptions

People assume

The model at the top of the table is the best one for us.

In fact

It is the best one at that paper. Whether it suits your material, your tone, your budget and your latency requirements is unrelated and unmeasured. Plenty of organisations run something several places down a table because it is faster, cheaper and better on the twenty cases they care about.

People assume

A few points of difference is meaningful.

In fact

Often it is within the noise of running the same test twice, and the ordering among close models changes with each release. Read a table as a grouping into rough tiers. Distinguishing first place from third is usually reading precision into a number that does not have it.

People assume

Benchmarks are objective measurements.

In fact

Somebody designed the test, chose what counts as correct and decided what to leave out, and vendors then choose which results to publish. All of that is legitimate and none of it is neutral. The number is real; what it is a number about is a set of decisions.

05

Questions

Should we choose a model by its benchmark scores?
Use them to narrow a long list to three or four, then decide on your own cases. Public scores are good at telling you which models are broadly in contention and poor at telling you which will suit your work, and the second question is the one you are actually asking when you buy.
What is contamination and why does it matter?
It is a benchmark's questions and answers ending up in the training material of a model later tested on them, which turns a test of ability into a test of memory. It cannot be reliably detected from outside, and it is the reason a striking score on a famous test deserves less weight than a modest score on a private one.
How do we build our own evaluation?
Collect twenty to fifty real cases including the awkward ones, write down what a good answer looks like for each, and have the person who does that work agree them. Run every candidate against all of them and count. Rerun whenever a model, a prompt or a document set changes. That is the whole method.
Why do the rankings keep changing?
Because models are released constantly, because close scores reorder on noise, and because new tests are introduced when old ones stop separating anything. A ranking is a snapshot of one paper on one day. Something you can rerun yourself is worth more precisely because it survives all that churn.
06

Key takeaways

  • A benchmark scores models on a fixed paper, which makes them comparable and not predictive of your work.
  • Published questions leak into training material, and it cannot be fully verified from outside.
  • Close scores reorder on noise, so read tables as tiers rather than as rankings.
  • What is measured is what can be marked automatically, which leaves out most of what business use depends on.
  • Twenty real cases with agreed answers is an afternoon's work and the only score that predicts your outcome.
08

Tools that use this

  • OpenRouter

    One interface for running your own cases across several models.

  • Claude

    A place to try your five real questions before reading any table.

  • ChatGPT

    The same question in two assistants, where your preference is the measurement.

Last checked July 2026

All glossary terms