Glossary
Retention policy
A retention policy states how long a provider keeps what you send it and what it sends back, where those copies live, and whether the period is something you can configure.
In plain terms
Every tool you send something to keeps it for a while. The policy is the statement of how long, which copies exist, and whether you have any say. With most software this is a dull paragraph nobody reads. With AI tools it matters more, because what gets sent is often a whole document rather than a form field, and because the answer differs sharply between the tier somebody signed up for on their own and the business tier an organisation would have bought.
Why it matters
Because it decides how long a decision you made once keeps mattering. Material sent to a tool in March is still somewhere in June if the period says so, and that is the window in which a question about it can be asked and has to be answered. It is also one of the few things that genuinely differs between account tiers of the same product, which is why it is the specific question worth asking rather than a general enquiry about security.
How it works
It covers both directions, and people usually only think about one. What you sent is the obvious half; what came back is a record too, and in a conversational tool the whole thread is typically retained as a unit rather than as separate messages. Asking about inputs alone leaves half the picture unexamined.
There are usually several copies with different clocks. The working copy that lets you scroll back through a conversation, backups, and anything kept for abuse monitoring or support are separate things with separate periods. A single headline number rarely describes all of them, which is why the useful question is what copies exist rather than how long you keep it.
Retention and training on your content are different questions that get conflated constantly. A provider can hold material for a period without ever using it to improve a model, and both commitments matter, but they are separate promises. Ask them separately, because a reassuring answer to one is frequently offered in response to the other.
Deletion means different things and the difference is worth pinning down. Removing something from your view, removing it from active systems, and removing it from backups are three distinct events that can be separated by a long interval. Where it matters, the question is when the last copy goes rather than when the item disappears from the interface.
The period is often configurable on administered tiers and fixed on individual ones, which is the practical reason this question keeps leading back to which account was used. A team on personal logins is accepting whatever the default is, usually without anybody having read it.
Where the copies go
Seen in the wild
Compare what a consumer tier and an administered organisational tier of the same assistant say about how long conversations are kept.
ChatGPTAsk what an enterprise search product retains about queries, given it reaches across systems where the material itself already lives.
GleanRun a model on your own hardware, where retention becomes a question about your storage rather than a vendor's policy.
Ollama
Common misconceptions
People assume
They said they do not train on our data, so retention is covered.
In fact
Those are separate commitments. Not training on content says nothing about how long it is kept, and a strong assurance on one is regularly offered when the question was about the other. Both are worth having in writing, and they belong in separate sentences.
People assume
Deleting the conversation removes it.
In fact
It usually removes it from your view and starts a clock on the rest. Active systems and backups have their own periods, and the interval between disappearing from the interface and disappearing entirely can be considerable. Where it matters, ask when the last copy goes.
Questions
- What is the single most useful question to ask a vendor?
- What copies exist and what the clock is on each, rather than how long you keep it. The headline number usually describes the working copy and not backups, support records or anything held for abuse monitoring. Asking about copies surfaces the parts a single figure conceals, without needing you to guess what they are.
- How does this differ from ordinary software retention?
- Mainly in what gets sent. Traditional tools receive structured fields; AI tools routinely receive whole documents, transcripts and correspondence because that is what makes them useful. The retention question therefore covers more sensitive material by default, which raises the stakes on a policy nobody previously read.
- Is a shorter period always better?
- Not always, and the trade is real. Shorter retention reduces exposure and also removes the history that makes features work, complicates support when something needs investigating, and can conflict with your own record-keeping needs. It is a decision to take deliberately per category of material rather than a dial to turn to minimum.
- Who should own this internally?
- Whoever owns data protection more broadly, with input from the people using the tool, because they know what actually gets sent. It tends to fall between roles: the risk function does not see the usage and the practitioners do not read the terms, which is how a default period ends up governing material nobody intended to be governed by it.
Key takeaways
- The policy covers what you sent and what came back, usually as one retained thread.
- Several copies exist with different clocks; a single headline number rarely covers them all.
- Retention and training on your content are separate promises, routinely conflated.
- Deletion from your view is not deletion from backups, and the gap can be long.
- The period is often configurable on administered tiers and fixed on individual ones.
Last checked July 2026