Glossary
Data loss prevention (DLP)
Software that watches for sensitive information leaving and blocks or flags it, which works in proportion to how recognisable that information is.
In plain terms
Somebody reading over your shoulder as you send things, stopping you when they spot a card number. They are good at card numbers because a card number looks like a card number. They are much less use on a document whose sensitivity is a matter of what it says rather than what it looks like.
Why it matters
Because it is the control organisations reach for after deciding that material leaving is a problem, and how much it helps depends almost entirely on what kind of material. Anything with a fixed shape is caught reliably. A strategy document, a customer conversation or a draft contract has no shape to match, and those are usually what people actually worry about.
How it works
Recognition is the whole of it, and it works by pattern. Card numbers, national insurance numbers and account identifiers have predictable formats, so matching them is dependable and the results are easy to defend. This is the case the category was built for and it remains what it does well.
Classification extends the reach and moves the difficulty rather than removing it. Marking documents as confidential lets the software act on the mark rather than the contents, which works exactly as well as the marking is maintained, and marking is manual work with no immediate reward.
Blocking and flagging are different products in practice. Blocking stops the action and interrupts somebody, which produces complaints when it is wrong; flagging produces a queue that somebody has to read, which produces nothing at all when nobody reads it. Most organisations discover which they wanted by trying the other.
Tuning is the real cost and it is ongoing rather than one-off. Rules that are too tight interrupt legitimate work until people find routes around them; rules that are too loose produce a queue nobody trusts. Neither failure announces itself, and the second is the commoner and quieter one.
The AI case is the hard one because there is no pattern to match. A person pasting a paragraph of a contract into a browser tab is sending text with no recognisable shape to a destination that looks like an ordinary website, and the only reliable signal is the destination itself rather than anything about the material.
How well it recognises what is leaving
Seen in the wild
Rules that recognise a card number reliably and say nothing about a strategy document pasted into a browser tab.
ChatGPTBlocking uploads to unapproved destinations, which is a decision about the destination rather than the contents.
ClaudeAn automation moving records where the content is ordinary and the volume is the only signal.
Make
Common misconceptions
People assume
It stops sensitive information leaving.
In fact
It stops information it can recognise. Structured identifiers are caught well; a document whose sensitivity is a matter of meaning has nothing to match against, and that is usually the material people had in mind when they bought it.
People assume
It is a purchase rather than an ongoing commitment.
In fact
The rules need continuous adjustment against real traffic. Too tight and people route around it, too loose and the queue stops being read, and both drift quietly because neither produces an error anybody sees.
Questions
- Will it stop people pasting into AI tools?
- Only by deciding about the destination rather than the content. A paragraph of a contract has no recognisable pattern, so the enforceable rule is which sites may receive uploads or text at all, which is a policy decision that happens to be implemented here rather than a detection capability.
- Should it block or only flag?
- Blocking produces immediate complaints when it is wrong, which at least means somebody is looking. Flagging produces a queue whose value depends entirely on whether anybody reads it, and an unread queue is indistinguishable from no control at all while looking like one.
- What is the realistic first thing to cover?
- Whatever has a fixed format and a regulator attached, because that is where recognition is dependable and the consequences of missing it are clearest. Extending the same approach to judgement-based material is a much larger commitment, and it depends entirely on document classification being maintained by people who gain nothing from doing it.
Key takeaways
- It works in proportion to how recognisable the material is.
- Structured identifiers are caught well; meaning-based sensitivity is not.
- For AI tools the enforceable rule is about destination, not content.
- An unread flagging queue looks like a control and is not one.
Last checked August 2026