Skip to content

Glossary

Bias

Bias is a system producing systematically worse outcomes for some groups than others, which is a property of the results it produces rather than of anyone's intent.

In plain terms

The system treats some people worse than others, consistently, in a way nobody chose. It is not an accusation about anybody's motives. It is a description of what comes out the other end when the same process is applied to different groups, and it can be present in a system every person who built it would describe as fair.

01

Why it matters

Because it is the failure with named consequences attached. Most ways a system disappoints produce a bad experience; this one produces different outcomes for people in protected categories, and that is a category of problem with regulators, lawyers and a public dimension. It is also the one a buyer inherits most completely, since the pattern arrives in the model and shows up in your decisions.

02

How it works

It enters through the material rather than through the code, which is why it is so hard to inspect. A model learns patterns from a body of text or images produced by people, and those materials carry the world's existing distributions. Nothing has to go wrong for the pattern to survive; the model is doing what it was built to do, which is reproduce what it saw.

Removing the obvious fields does not remove it, and this is the part that most often surprises. Other information stands in: where somebody lives, where they studied, how they phrase things, what they have done before. A system given none of the protected attributes can still reconstruct them well enough to produce the same divergence, which is why deleting a column is not a treatment.

It is measured on outcomes rather than found in the system, which changes what an assessment looks like. The question is whether results differ across groups for the same input, and answering it requires data about those groups and a decision about which difference matters. That is a judgement, not a calculation, and it is the reason two honest assessments of the same system can disagree.

It is a property of a deployment, not a fixed attribute of a model. The same model used for two purposes can be unproblematic in one and not the other, because what counts as a worse outcome depends entirely on what the output decides. This is why a vendor cannot certify the absence of it, and why an agent platform in this guide says plainly that it does not substitute for compliance review of how its agents behave.

The most consequential uses are the ones where a system sorts people, which is where the term stops being abstract. Hiring, lending, admissions, allocation and anything that ranks candidates share a shape: the output orders human beings, at a volume no individual reviewer could match, applying whatever pattern it learned consistently to every one of them. Consistency is the feature, and it is also what turns a small divergence into a systematic one.

Where people look for it, and where it is

Where people look for it, and where it isThis is the most common way the subject is mishandled, and it is mishandled in good faith. A team asked to check for bias does what can be done with what is available: it inspects the inputs, confirms no protected attribute is present, reviews the prompt for loaded language, and reports back honestly that nothing was found. Every step is reasonable and the conclusion does not follow, because the property being looked for does not live in any of those places. It lives in the comparison between what happened to one group and what happened to another, which requires knowing who was in each group. Most organisations do not hold that information, and cannot casually collect it, so the assessment that would actually answer the question is the expensive one and the reassuring one is free. Naming that gap is more useful than any individual finding, because a team that knows it has run the cheap check will not mistake the result for the answer.Where attention goesThe protected fields in thedata.The wording of the prompt.Whether anyone intended harm.All inspectable, all cheap tocheck.Where it actually showsOutcomes compared acrossgroups.Fields that stand in for theremoved ones.The specific decision beingmade.Needs data most organisationslack.The left column is checkable inan afternoon and answers adifferent question. Thatasymmetry is why so many systemshave been reviewed for this andnever assessed for it.
This is the most common way the subject is mishandled, and it is mishandled in good faith. A team asked to check for bias does what can be done with what is available: it inspects the inputs, confirms no protected attribute is present, reviews the prompt for loaded language, and reports back honestly that nothing was found. Every step is reasonable and the conclusion does not follow, because the property being looked for does not live in any of those places. It lives in the comparison between what happened to one group and what happened to another, which requires knowing who was in each group. Most organisations do not hold that information, and cannot casually collect it, so the assessment that would actually answer the question is the expensive one and the reassuring one is free. Naming that gap is more useful than any individual finding, because a team that knows it has run the cheap check will not mistake the result for the answer.
03

Seen in the wild

  • An agent platform stating that it does not substitute for compliance review of how its agents behave, which leaves the assessment with the buyer.

    Sierra
  • Matched candidates surfaced ranked for review, which is the shape where a system orders people at volume rather than answering a question.

    LinkedIn Recruiter
  • A generator whose training data is licensed and public-domain rather than scraped, which changes what distribution the output reproduces.

    Adobe Firefly
04

Common misconceptions

People assume

A model can be certified as unbiased.

In fact

Not in general, because it is a property of an outcome on a population rather than of the model alone. The same model can be unproblematic for one use and not for another, so what can be assessed is a specific deployment, with its own data, on the groups it actually affects.

People assume

Removing protected attributes fixes it.

In fact

Other fields stand in for them. Postcode, education, phrasing and history carry enough signal to reconstruct the pattern, so a system that has never seen a protected attribute can reproduce the same divergence. This is why an assessment has to look at outcomes rather than at inputs.

05

Telling them apart

Bias vs Explainability

Bias

Whether outcomes differ across groups.

Explainability

Whether anyone can say why a particular outcome happened.

You can have either without the other. An explainable system can be reliably unfair, and a fair one can be entirely opaque.

06

Questions

Is this only a concern for systems that decide about people?
Those are where the consequences are named, but the pattern is broader. Drafting tools carrying assumptions about who does which job, and generators producing narrow depictions of ordinary roles, shape material that goes to customers and staff. The difference is that one produces a decision somebody can appeal and the other produces an impression nobody logs.
Whose responsibility is it when the model came from a vendor?
The deployment is yours, and vendors increasingly say so in their own documentation. The pattern may originate upstream, and the decision to apply it here, to these people, is made by the buyer. That split is uncomfortable and it is the one reflected in how the obligations actually land.
Can it be tested before deployment?
Partly, and the limit is worth knowing in advance. Testing needs data about the groups in question and agreement on which difference counts, and organisations often have neither to hand. That is a real obstacle rather than an excuse, but it does mean the honest answer before launch is usually a range rather than a verdict.
07

Key takeaways

  • A property of outcomes on a population, not of anyone's intent or of a model alone.
  • It arrives through the training material, which is the model working as designed.
  • Removing protected fields does not remove it; other fields stand in for them.
  • It cannot be certified in general, only assessed for a specific deployment.
  • The sharpest cases are systems that sort people, where consistency scales the divergence.
09

Tools that use this

  • Sierra

    Compliance review of agent behaviour left with the buyer, not the platform.

  • LinkedIn Recruiter

    Candidates surfaced ranked for review, the shape where a system orders people.

  • Adobe Firefly

    Licensed training data, which changes the distribution reproduced.

Last checked July 2026

All glossary terms