Glossary
Anonymisation
Altering data so that no individual can be identified from it again, which is a much higher bar than removing the obvious identifiers.
In plain terms
Changing the data so nobody can work out who it is about, ever, by any route. The word doing the work is ever: a version that is hard to trace back is not anonymous, it is difficult, and the two are treated completely differently.
Why it matters
Because it is the only step that genuinely removes material from scope, which makes it enormously attractive and correspondingly over-claimed. An organisation that has achieved it has fewer obligations about that data; one that believes it has achieved it and has not is operating with no protections and no obligations either.
How it works
The test is identifiability by any reasonably available means, which is why removing names does not settle it. What matters is whether somebody could get back to an individual using what they have plus what they could obtain, and that is a question about the wider world rather than about the file.
Combinations identify people even when individual fields do not. A handful of ordinary attributes together can narrow a population to one person, so a dataset where every column looks harmless can still be identifiable, and the assessment has to consider the columns as a set.
It is one-way by definition, which is what separates it from every technique that keeps a route back. If a mapping exists anywhere that could restore identity, the data has been protected rather than anonymised, and it remains within scope for as long as that mapping exists.
The bar tends to be higher than teams expect and regulator views on where it sits differ. That is why an internal declaration that a dataset is anonymous is worth less than it seems, particularly when the declaration was made by the people who wanted to use the data.
For AI work the attraction is obvious and the trap is specific. Anonymised material can be used far more freely, so the incentive to declare success is strong, and the datasets richest enough to be useful for training or analysis are exactly the ones where combinations of fields make identification easiest.
What the claim requires
Seen in the wild
Stripping names from support conversations, where the remaining detail still identifies the customer.
ChatGPTA dataset prepared for analysis where several ordinary columns together narrow to individuals.
Julius AIMaterial described internally as anonymised by the team that wanted to use it.
Power BI Copilot
Common misconceptions
People assume
Removing names and identifiers makes data anonymous.
In fact
It removes the obvious route and rarely all of them. Combinations of ordinary attributes narrow a population quickly, so the test is whether anybody could get back to an individual by any reasonably available means rather than whether the name column is gone.
People assume
If it is difficult to identify somebody, that is enough.
In fact
Difficulty and impossibility are treated very differently. Where a route back exists at all, the data is protected rather than anonymised, and it stays within scope along with every obligation that carries.
Telling them apart
Anonymisation vs Pseudonymisation
Anonymisation
No route back exists; the data leaves scope entirely.
A route back exists and is held separately; the data stays in scope.
Ask whether anything anywhere could restore identity. If yes, it is the second one whatever it is called.
Questions
- How do we know whether we have actually achieved it?
- By asking whether anybody could get back to an individual using what they hold plus what they could reasonably obtain, considering the columns together rather than one at a time. That is a genuine assessment, and it is worth more when the people doing it are not the people wanting to use the data.
- Why is the bar so high?
- Because clearing it removes the data from scope entirely, which is a substantial consequence. A test that was easy to pass would let any organisation exempt itself by declaration, so the standard is set at a level that requires the claim to be demonstrable.
- Is it worth attempting for AI work?
- Where it can genuinely be achieved it is the strongest position available, and the datasets rich enough to be useful are the hardest to anonymise. That tension is real, so it is worth attempting with the expectation that the honest answer may be that it was not reached.
Key takeaways
- The test is identifiability by any reasonably available means.
- Combinations of ordinary fields identify people that single fields do not.
- If a route back exists anywhere, it is not this.
- The richest datasets are the hardest to anonymise, which is the AI trap.
Tools that use this
- ChatGPT
Names stripped from conversations that still identify the customer.
- Julius AI
Ordinary columns that together narrow a dataset to individuals.
- Power BI Copilot
Declared anonymous by the team that wanted to use it.
Last checked August 2026