Glossary
Benchmark
A benchmark is a standard test scoring models against each other on a fixed set of questions, which makes it useful for drawing up a shortlist and unreliable as a prediction about your own work.
In plain terms
It is an exam with a published paper. Everybody sits the same questions and the marks get compared. That works well for spotting who is hopeless and works badly for predicting who will be good at your particular job, partly because your job is not on the paper and partly because by now everybody has seen the paper.
Why it matters
Scores drive purchasing decisions, press coverage and a great deal of internal argument, and the gap between scoring well and working for you is where the money goes. The useful reframing is that you are allowed to write your own exam. Twenty real cases from your own work, with answers somebody competent agrees are good, will tell you more in an afternoon than every published table put together, and it keeps telling you when the models change underneath.
How it works
A fixed set of questions with known answers is run identically across models and the results are scored, producing a number that can be put in a table. The value of the arrangement is comparability: whatever else is true, every model faced the same paper, which is more than can be said for most claims in this market.
Published tests leak into training material, and this is the central problem rather than an edge case. A model trained after a benchmark became public may have encountered its questions and answers, which turns a test of ability into a test of recall. Nobody can fully verify whether it happened, which is why a high score on a well-known test is weaker evidence than it looks.
Scores are frequently close and the ordering is less stable than the reporting suggests. Small differences fall within the variation you would get from running the same test twice, and a lead that changes hands between two releases was probably never a lead. Treat a table as a rough grouping into tiers rather than as a ranking.
What gets measured is narrow, because it has to be automatically checkable. Exam-style questions with definite answers work well. Tone, judgement, following a house style, knowing when to decline and being pleasant to work with over an hour are largely unmeasured, and for a great many business uses those are the whole of what you are buying.
Your own set is the answer and it is less work than it sounds. Twenty to fifty real cases from your actual work, with agreed good answers, run whenever a model or a prompt changes. It is the same discipline that makes prompt work repeatable, and it is the only score that has any bearing on whether the thing will help you.
Two exams, and which one is about you
Seen in the wild
Run twenty of your own real cases across several models through one routing interface, which is the comparison that actually predicts something about your work.
OpenRouterTake the five questions you genuinely need answered well and put them to an assistant before reading a single published score.
ClaudeAsk the same question in two assistants and notice how often your own preference disagrees with the published ranking.
ChatGPT
Common misconceptions
People assume
The model at the top of the table is the best one for us.
In fact
It is the best one at that paper. Whether it suits your material, your tone, your budget and your latency requirements is unrelated and unmeasured. Plenty of organisations run something several places down a table because it is faster, cheaper and better on the twenty cases they care about.
People assume
A few points of difference is meaningful.
In fact
Often it is within the noise of running the same test twice, and the ordering among close models changes with each release. Read a table as a grouping into rough tiers. Distinguishing first place from third is usually reading precision into a number that does not have it.
People assume
Benchmarks are objective measurements.
In fact
Somebody designed the test, chose what counts as correct and decided what to leave out, and vendors then choose which results to publish. All of that is legitimate and none of it is neutral. The number is real; what it is a number about is a set of decisions.
Questions
- Should we choose a model by its benchmark scores?
- Use them to narrow a long list to three or four, then decide on your own cases. Public scores are good at telling you which models are broadly in contention and poor at telling you which will suit your work, and the second question is the one you are actually asking when you buy.
- What is contamination and why does it matter?
- It is a benchmark's questions and answers ending up in the training material of a model later tested on them, which turns a test of ability into a test of memory. It cannot be reliably detected from outside, and it is the reason a striking score on a famous test deserves less weight than a modest score on a private one.
- How do we build our own evaluation?
- Collect twenty to fifty real cases including the awkward ones, write down what a good answer looks like for each, and have the person who does that work agree them. Run every candidate against all of them and count. Rerun whenever a model, a prompt or a document set changes. That is the whole method.
- Why do the rankings keep changing?
- Because models are released constantly, because close scores reorder on noise, and because new tests are introduced when old ones stop separating anything. A ranking is a snapshot of one paper on one day. Something you can rerun yourself is worth more precisely because it survives all that churn.
Key takeaways
- A benchmark scores models on a fixed paper, which makes them comparable and not predictive of your work.
- Published questions leak into training material, and it cannot be fully verified from outside.
- Close scores reorder on noise, so read tables as tiers rather than as rankings.
- What is measured is what can be marked automatically, which leaves out most of what business use depends on.
- Twenty real cases with agreed answers is an afternoon's work and the only score that predicts your outcome.
Tools that use this
- OpenRouter
One interface for running your own cases across several models.
- Claude
A place to try your five real questions before reading any table.
- ChatGPT
The same question in two assistants, where your preference is the measurement.
Last checked July 2026