Glossary
Pseudonymisation
Replacing identifying details with references that can be reversed using information held separately, which reduces risk without changing scope.
In plain terms
Swapping names for codes and keeping the codebook somewhere else. Anybody who gets the data without the codebook has much less than they would have had, and the codebook exists, which is why the data is still about people.
Why it matters
Because it is the measure organisations actually achieve, as opposed to the one they claim. It genuinely reduces what a breach exposes and what most staff can see, and it leaves every obligation in place, so an organisation treating it as an exemption has taken a real protection and drawn a false conclusion from it.
How it works
The separation is the whole mechanism. Identifying details are replaced with references, and the information needed to reverse that is held apart under its own controls, so the protection is exactly as strong as that separation is maintained.
It stays personal data because the route back exists. That is not a technicality: the individual can still be reached by whoever holds both parts, so the material continues to concern an identifiable person and the rules continue to apply to it.
Its real value is narrowing exposure rather than removing it. A copy that leaks without the additional information is much less damaging than one that leaks with it, and most people inside an organisation can work with the pseudonymised version without ever needing the other half.
It fails quietly when the separation erodes. The two halves drifting into the same system, the same backup or the same person's access removes the protection while everything continues to look correct, and nothing about the data itself reveals that this has happened.
For AI work it is frequently the practical answer. Analysis, classification and search often work perfectly well on pseudonymised material, so the tool gets what it needs while the identifying half stays out of the deployment entirely, which is a better trade than most of the alternatives.
The question worth asking a vendor is who holds the mapping and where. Where a tool performs the substitution itself and stores both halves, the separation exists in the product's design rather than between two organisations, and that is a materially weaker arrangement than one where the identifying half never leaves your systems at all.
What changes and what does not
Seen in the wild
Common misconceptions
People assume
It takes the data out of scope.
In fact
It does not, because a route back exists and somebody holds it. The obligations continue in full, and an organisation treating the measure as an exemption has taken a genuine protection and drawn the wrong conclusion from it.
People assume
It is a weaker version of anonymisation.
In fact
It is a different thing with a different purpose. Anonymisation removes data from scope and is rarely achieved; this reduces exposure while keeping the data usable and reversible, which is frequently what the situation actually needs.
Questions
- If it does not remove obligations, why do it?
- Because it reduces what a breach exposes and what most people can see, which are real improvements independent of scope. It also lets analysis proceed on material that never needs to identify anybody, so the identifying half stays out of the tool entirely.
- How does it fail in practice?
- By the separation eroding rather than by anybody breaking it. The two halves ending up in the same system, the same backup or within one person's access removes the protection while everything continues to look exactly as it did, and the data itself gives no sign.
- Is it enough for an AI deployment?
- It is frequently the right measure, and it is not a substitute for the other questions. The material remains personal data, so the ground for processing it, how long it is kept and where it goes all still need answers, which pseudonymisation does not provide.
Key takeaways
- The route back exists, so the data stays personal data.
- The protection is exactly as strong as the separation is maintained.
- It fails by erosion, invisibly, with nothing about the data changing.
- For AI work it is often the practical answer rather than a compromise.
Last checked August 2026