Skip to content

Glossary

Anonymisation

Altering data so that no individual can be identified from it again, which is a much higher bar than removing the obvious identifiers.

In plain terms

Changing the data so nobody can work out who it is about, ever, by any route. The word doing the work is ever: a version that is hard to trace back is not anonymous, it is difficult, and the two are treated completely differently.

01

Why it matters

Because it is the only step that genuinely removes material from scope, which makes it enormously attractive and correspondingly over-claimed. An organisation that has achieved it has fewer obligations about that data; one that believes it has achieved it and has not is operating with no protections and no obligations either.

02

How it works

The test is identifiability by any reasonably available means, which is why removing names does not settle it. What matters is whether somebody could get back to an individual using what they have plus what they could obtain, and that is a question about the wider world rather than about the file.

Combinations identify people even when individual fields do not. A handful of ordinary attributes together can narrow a population to one person, so a dataset where every column looks harmless can still be identifiable, and the assessment has to consider the columns as a set.

It is one-way by definition, which is what separates it from every technique that keeps a route back. If a mapping exists anywhere that could restore identity, the data has been protected rather than anonymised, and it remains within scope for as long as that mapping exists.

The bar tends to be higher than teams expect and regulator views on where it sits differ. That is why an internal declaration that a dataset is anonymous is worth less than it seems, particularly when the declaration was made by the people who wanted to use the data.

For AI work the attraction is obvious and the trap is specific. Anonymised material can be used far more freely, so the incentive to declare success is strong, and the datasets richest enough to be useful for training or analysis are exactly the ones where combinations of fields make identification easiest.

What the claim requires

What the claim requiresWhat makes this gap durable is that the left-hand column is genuinely reassuring to look at. A file with no names in it reads as anonymous, and the person who prepared it did real work and can see the result of that work, while the right-hand column is invisible: it asks about routes that exist elsewhere, combinations somebody else might exploit and mappings held by another team. None of that is apparent from the file. The incentive runs the same way, because the reward for the claim is substantial and immediate and the cost of being wrong is deferred and probabilistic, so an organisation is deciding under exactly the conditions where optimism wins. For AI work the tension is sharpest and worth stating without a resolution, because there may not be one: the datasets valuable enough to justify the effort are detailed, and detail is what makes people identifiable. An honest attempt frequently ends with the conclusion that anonymisation was not achieved, and that conclusion is a useful output rather than a failed exercise.Frequently doneThe name column is removed.Identifiers are replaced.It looks anonymous to a reader.What the test asksCould anyone get back by anyroute?Do the columns identify incombination?Does a mapping surviveanywhere?The left-hand column is what thephrase is usually used todescribe and the right-handcolumn is what it means. Adataset satisfying only the leftis protected, which is worthsomething, and is still inscope.
What makes this gap durable is that the left-hand column is genuinely reassuring to look at. A file with no names in it reads as anonymous, and the person who prepared it did real work and can see the result of that work, while the right-hand column is invisible: it asks about routes that exist elsewhere, combinations somebody else might exploit and mappings held by another team. None of that is apparent from the file. The incentive runs the same way, because the reward for the claim is substantial and immediate and the cost of being wrong is deferred and probabilistic, so an organisation is deciding under exactly the conditions where optimism wins. For AI work the tension is sharpest and worth stating without a resolution, because there may not be one: the datasets valuable enough to justify the effort are detailed, and detail is what makes people identifiable. An honest attempt frequently ends with the conclusion that anonymisation was not achieved, and that conclusion is a useful output rather than a failed exercise.
03

Seen in the wild

  • Stripping names from support conversations, where the remaining detail still identifies the customer.

    ChatGPT
  • A dataset prepared for analysis where several ordinary columns together narrow to individuals.

    Julius AI
  • Material described internally as anonymised by the team that wanted to use it.

    Power BI Copilot
04

Common misconceptions

People assume

Removing names and identifiers makes data anonymous.

In fact

It removes the obvious route and rarely all of them. Combinations of ordinary attributes narrow a population quickly, so the test is whether anybody could get back to an individual by any reasonably available means rather than whether the name column is gone.

People assume

If it is difficult to identify somebody, that is enough.

In fact

Difficulty and impossibility are treated very differently. Where a route back exists at all, the data is protected rather than anonymised, and it stays within scope along with every obligation that carries.

05

Telling them apart

Anonymisation vs Pseudonymisation

Anonymisation

No route back exists; the data leaves scope entirely.

Pseudonymisation

A route back exists and is held separately; the data stays in scope.

Ask whether anything anywhere could restore identity. If yes, it is the second one whatever it is called.

06

Questions

How do we know whether we have actually achieved it?
By asking whether anybody could get back to an individual using what they hold plus what they could reasonably obtain, considering the columns together rather than one at a time. That is a genuine assessment, and it is worth more when the people doing it are not the people wanting to use the data.
Why is the bar so high?
Because clearing it removes the data from scope entirely, which is a substantial consequence. A test that was easy to pass would let any organisation exempt itself by declaration, so the standard is set at a level that requires the claim to be demonstrable.
Is it worth attempting for AI work?
Where it can genuinely be achieved it is the strongest position available, and the datasets rich enough to be useful are the hardest to anonymise. That tension is real, so it is worth attempting with the expectation that the honest answer may be that it was not reached.
07

Key takeaways

  • The test is identifiability by any reasonably available means.
  • Combinations of ordinary fields identify people that single fields do not.
  • If a route back exists anywhere, it is not this.
  • The richest datasets are the hardest to anonymise, which is the AI trap.
09

Tools that use this

  • ChatGPT

    Names stripped from conversations that still identify the customer.

  • Julius AI

    Ordinary columns that together narrow a dataset to individuals.

  • Power BI Copilot

    Declared anonymous by the team that wanted to use it.

Last checked August 2026

All glossary terms