Glossary
Drift detection
Drift detection is noticing that results have got worse because something outside your control moved, while nothing in your own estate was changed.
In plain terms
You changed nothing and the answers are worse. Nobody deployed anything. What moved was outside your control: the questions people ask, the documents it reads, the vocabulary of the business, and sometimes the model the product runs on. It was fitted to a world, the world carried on, and the fit loosened without anything failing.
Why it matters
Because it is the failure with no owner. Every change-related problem has somebody who made the change, and this one has nobody, so it does not surface in any review and belongs to no team. It is also the most likely reason a system that was genuinely good at launch is quietly disappointing a year later, which is a pattern most organisations have experienced without naming.
How it works
Four different things move, and separating them decides what to do. The inputs change, meaning people ask about new things or in new ways. The source material changes, so what the system reads is now stale or reorganised. The world changes, and yesterday's correct answer is no longer correct. Or the product quietly moves to a different model underneath you, which you were not told about. The symptom is identical in all four and the remedy is not.
The stale-source case is the one most buyers actually meet, because it needs no machine learning to occur. Enterprise search in this guide states it directly: answer quality inherits the estate's hygiene, and stale documents produce well-cited stale answers. Nothing about the system degraded; it is faithfully reporting material that stopped being true.
Detection is comparison against a baseline, which is why it depends entirely on monitoring having been in place beforehand. Drift is a movement, and a movement cannot be observed from a single reading. An organisation noticing this for the first time usually cannot say when it began, which is the practical cost of starting to measure after the complaint rather than before.
It is gradual, and gradual is what makes it hard to see from inside. A sharp drop gets investigated because it looks like a fault. A slow decline gets absorbed: people adjust how they phrase things, stop asking about the areas that answer badly, and recalibrate what they expect. By the time it is raised, the same users have usually been compensating for months.
What moved, when nothing was deployed
Seen in the wild
Enterprise search whose answer quality inherits the estate's hygiene, where stale documents produce well-cited stale answers.
GleanA workspace assistant answering from your own content, where the answers track whatever the workspace currently says rather than what is true.
Notion AIAutonomy earned rather than assumed, with supervision preceding trust, which is the posture that keeps a slow decline visible.
Relevance AI
Common misconceptions
People assume
The model degraded.
In fact
Models do not decay on their own, so the one you were using did not get worse. Two different things are being confused here. A model can be swapped underneath you, because a purchased tool can change what it runs on without telling you, and that is a real cause worth checking. What it is not is decay, and it is not the only candidate: the questions, the source material and the facts move too. Sources are worth checking first because they are the cheapest and commonest answer, not because a change of model is unlikely.
People assume
It will be obvious when it happens.
In fact
It is gradual by nature, and people adapt to it faster than they report it. Users rephrase, avoid the topics that answer badly, and lower their expectations, all of which suppress exactly the signal that would have surfaced the problem. That adaptation is why the complaint arrives long after the decline began.
Telling them apart
Drift detection vs Model monitoring
Drift detection
One cause: something outside your control moved.
The practice of watching output at all, whatever the cause.
You cannot detect drift without monitoring, because drift is a change and a change needs two measurements.
Questions
- How is this noticed at all?
- By comparing against something recorded earlier, which is the whole method. Without a baseline there is only today's reading and an impression that things used to be better, and an impression cannot distinguish a genuine decline from a run of harder questions or from expectations having risen.
- What is the usual cause for a bought tool?
- Stale source material, more often than anything else. A system answering from your documents is only as current as the documents, and organisations refresh their content far less often than they assume. That case needs no retraining to explain and no retraining to fix.
- Does this apply to assistants as well as to trained models?
- Yes, and it is arguably more relevant there, because the material behind an assistant changes constantly while the model behind it changes on somebody else's schedule. Both move, neither is under your control, and only the combination is visible in what a user experiences.
Key takeaways
- You changed nothing; the inputs, the sources, the world or the model underneath moved.
- Stale sources are the case most buyers actually meet, and need no retraining to explain.
- It is a movement, so detecting it requires a baseline recorded beforehand.
- Gradual by nature, and users adapt to it faster than they report it.
Tools that use this
- Glean
Answer quality inheriting the estate's hygiene, stale documents included.
- Notion AI
Answers tracking whatever the workspace currently says.
- Relevance AI
Supervision preceding trust, which keeps a slow decline visible.
Last checked July 2026