Skip to content

Glossary

Pseudonymisation

Replacing identifying details with references that can be reversed using information held separately, which reduces risk without changing scope.

In plain terms

Swapping names for codes and keeping the codebook somewhere else. Anybody who gets the data without the codebook has much less than they would have had, and the codebook exists, which is why the data is still about people.

01

Why it matters

Because it is the measure organisations actually achieve, as opposed to the one they claim. It genuinely reduces what a breach exposes and what most staff can see, and it leaves every obligation in place, so an organisation treating it as an exemption has taken a real protection and drawn a false conclusion from it.

02

How it works

The separation is the whole mechanism. Identifying details are replaced with references, and the information needed to reverse that is held apart under its own controls, so the protection is exactly as strong as that separation is maintained.

It stays personal data because the route back exists. That is not a technicality: the individual can still be reached by whoever holds both parts, so the material continues to concern an identifiable person and the rules continue to apply to it.

Its real value is narrowing exposure rather than removing it. A copy that leaks without the additional information is much less damaging than one that leaks with it, and most people inside an organisation can work with the pseudonymised version without ever needing the other half.

It fails quietly when the separation erodes. The two halves drifting into the same system, the same backup or the same person's access removes the protection while everything continues to look correct, and nothing about the data itself reveals that this has happened.

For AI work it is frequently the practical answer. Analysis, classification and search often work perfectly well on pseudonymised material, so the tool gets what it needs while the identifying half stays out of the deployment entirely, which is a better trade than most of the alternatives.

The question worth asking a vendor is who holds the mapping and where. Where a tool performs the substitution itself and stores both halves, the separation exists in the product's design rather than between two organisations, and that is a materially weaker arrangement than one where the identifying half never leaves your systems at all.

What changes and what does not

What changes and what does notThe reason this misunderstanding is so common is that the measure feels like a transformation of the data and is really a transformation of who can reach what. Somebody looking at a pseudonymised file sees codes where names were, which resembles the thing that would take it out of scope, and the difference between the two lives entirely in whether a mapping exists somewhere else. That is a fact about the organisation rather than about the file, which is why it cannot be seen by inspecting the data and why teams reach the wrong conclusion in good faith. The practical consequence is worth stating plainly, because it is the opposite of discouraging: pseudonymisation is usually the right thing to do, it is achievable where anonymisation is not, and it makes AI deployments meaningfully safer by keeping the identifying half out of the tool. What it does not do is answer any of the questions in the right-hand column, and those need answering separately rather than being considered settled by the technique.Genuinely improvedWhat a leaked copy exposes.What most staff can see.What a tool needs to hold.Entirely unchangedWhether the rules apply.The ground you rely on.How long it may be kept.The left-hand column is worthhaving and is why the measure isrecommended. The right-handcolumn is what organisationsassume they have bought, and notechnique applied to the datacan deliver it while a routeback exists.
The reason this misunderstanding is so common is that the measure feels like a transformation of the data and is really a transformation of who can reach what. Somebody looking at a pseudonymised file sees codes where names were, which resembles the thing that would take it out of scope, and the difference between the two lives entirely in whether a mapping exists somewhere else. That is a fact about the organisation rather than about the file, which is why it cannot be seen by inspecting the data and why teams reach the wrong conclusion in good faith. The practical consequence is worth stating plainly, because it is the opposite of discouraging: pseudonymisation is usually the right thing to do, it is achievable where anonymisation is not, and it makes AI deployments meaningfully safer by keeping the identifying half out of the tool. What it does not do is answer any of the questions in the right-hand column, and those need answering separately rather than being considered settled by the technique.
03

Seen in the wild

  • Replacing customer identifiers before conversations are analysed, keeping the mapping outside the tool.

    ChatGPT
  • A dataset prepared for analysis where the reference table sits in a different system entirely.

    Julius AI
  • An automation working on coded records without ever needing the identifying half.

    Make
04

Common misconceptions

People assume

It takes the data out of scope.

In fact

It does not, because a route back exists and somebody holds it. The obligations continue in full, and an organisation treating the measure as an exemption has taken a genuine protection and drawn the wrong conclusion from it.

People assume

It is a weaker version of anonymisation.

In fact

It is a different thing with a different purpose. Anonymisation removes data from scope and is rarely achieved; this reduces exposure while keeping the data usable and reversible, which is frequently what the situation actually needs.

05

Questions

If it does not remove obligations, why do it?
Because it reduces what a breach exposes and what most people can see, which are real improvements independent of scope. It also lets analysis proceed on material that never needs to identify anybody, so the identifying half stays out of the tool entirely.
How does it fail in practice?
By the separation eroding rather than by anybody breaking it. The two halves ending up in the same system, the same backup or within one person's access removes the protection while everything continues to look exactly as it did, and the data itself gives no sign.
Is it enough for an AI deployment?
It is frequently the right measure, and it is not a substitute for the other questions. The material remains personal data, so the ground for processing it, how long it is kept and where it goes all still need answers, which pseudonymisation does not provide.
06

Key takeaways

  • The route back exists, so the data stays personal data.
  • The protection is exactly as strong as the separation is maintained.
  • It fails by erosion, invisibly, with nothing about the data changing.
  • For AI work it is often the practical answer rather than a compromise.
08

Tools that use this

  • ChatGPT

    Replacing identifiers before analysis, mapping kept outside the tool.

  • Julius AI

    A reference table held in a different system from the dataset.

  • Make

    Working on coded records without needing the identifying half.

Last checked August 2026

All glossary terms