Skip to content

Glossary

Drift detection

Drift detection is noticing that results have got worse because something outside your control moved, while nothing in your own estate was changed.

In plain terms

You changed nothing and the answers are worse. Nobody deployed anything. What moved was outside your control: the questions people ask, the documents it reads, the vocabulary of the business, and sometimes the model the product runs on. It was fitted to a world, the world carried on, and the fit loosened without anything failing.

01

Why it matters

Because it is the failure with no owner. Every change-related problem has somebody who made the change, and this one has nobody, so it does not surface in any review and belongs to no team. It is also the most likely reason a system that was genuinely good at launch is quietly disappointing a year later, which is a pattern most organisations have experienced without naming.

02

How it works

Four different things move, and separating them decides what to do. The inputs change, meaning people ask about new things or in new ways. The source material changes, so what the system reads is now stale or reorganised. The world changes, and yesterday's correct answer is no longer correct. Or the product quietly moves to a different model underneath you, which you were not told about. The symptom is identical in all four and the remedy is not.

The stale-source case is the one most buyers actually meet, because it needs no machine learning to occur. Enterprise search in this guide states it directly: answer quality inherits the estate's hygiene, and stale documents produce well-cited stale answers. Nothing about the system degraded; it is faithfully reporting material that stopped being true.

Detection is comparison against a baseline, which is why it depends entirely on monitoring having been in place beforehand. Drift is a movement, and a movement cannot be observed from a single reading. An organisation noticing this for the first time usually cannot say when it began, which is the practical cost of starting to measure after the complaint rather than before.

It is gradual, and gradual is what makes it hard to see from inside. A sharp drop gets investigated because it looks like a fault. A slow decline gets absorbed: people adjust how they phrase things, stop asking about the areas that answer badly, and recalibrate what they expect. By the time it is raised, the same users have usually been compensating for months.

What moved, when nothing was deployed

What moved, when nothing was deployedOrdering the causes this way is useful because the symptom gives no clue which one you have, and the instinct is to reach for the right-hand end first. Somebody reports that the answers have got worse, and the conversation turns immediately to the model, to retraining, to whether a different provider would do better. The left-hand end is far more likely and far cheaper, and it is checkable in an afternoon: take a handful of the answers people are unhappy with and look at the documents behind them. A striking proportion of the time the system has faithfully reported something that was true when it was written and has not been touched since. Starting at the left is not merely cheaper, it also produces the evidence needed to justify looking further right if the answer is not there. Note the fourth mark, which is not a fault in anything: a purchased tool can change what it runs on without telling you, so it belongs on this picture even though nobody on your side moved.EASIEST TO FIXHARDESTStale sourcesThe documentsstopped beingtrue. Refreshthem.New questionsPeople askabout thingsnever covered.ChangedvocabularyThe businessrenamed whatit does.A differentmodelunderneathThe productmoved, and didnot say so.The worldmovedYesterday'scorrect answeris now wrong.
Ordering the causes this way is useful because the symptom gives no clue which one you have, and the instinct is to reach for the right-hand end first. Somebody reports that the answers have got worse, and the conversation turns immediately to the model, to retraining, to whether a different provider would do better. The left-hand end is far more likely and far cheaper, and it is checkable in an afternoon: take a handful of the answers people are unhappy with and look at the documents behind them. A striking proportion of the time the system has faithfully reported something that was true when it was written and has not been touched since. Starting at the left is not merely cheaper, it also produces the evidence needed to justify looking further right if the answer is not there. Note the fourth mark, which is not a fault in anything: a purchased tool can change what it runs on without telling you, so it belongs on this picture even though nobody on your side moved.
03

Seen in the wild

  • Enterprise search whose answer quality inherits the estate's hygiene, where stale documents produce well-cited stale answers.

    Glean
  • A workspace assistant answering from your own content, where the answers track whatever the workspace currently says rather than what is true.

    Notion AI
  • Autonomy earned rather than assumed, with supervision preceding trust, which is the posture that keeps a slow decline visible.

    Relevance AI
04

Common misconceptions

People assume

The model degraded.

In fact

Models do not decay on their own, so the one you were using did not get worse. Two different things are being confused here. A model can be swapped underneath you, because a purchased tool can change what it runs on without telling you, and that is a real cause worth checking. What it is not is decay, and it is not the only candidate: the questions, the source material and the facts move too. Sources are worth checking first because they are the cheapest and commonest answer, not because a change of model is unlikely.

People assume

It will be obvious when it happens.

In fact

It is gradual by nature, and people adapt to it faster than they report it. Users rephrase, avoid the topics that answer badly, and lower their expectations, all of which suppress exactly the signal that would have surfaced the problem. That adaptation is why the complaint arrives long after the decline began.

05

Telling them apart

Drift detection vs Model monitoring

Drift detection

One cause: something outside your control moved.

Model monitoring

The practice of watching output at all, whatever the cause.

You cannot detect drift without monitoring, because drift is a change and a change needs two measurements.

06

Questions

How is this noticed at all?
By comparing against something recorded earlier, which is the whole method. Without a baseline there is only today's reading and an impression that things used to be better, and an impression cannot distinguish a genuine decline from a run of harder questions or from expectations having risen.
What is the usual cause for a bought tool?
Stale source material, more often than anything else. A system answering from your documents is only as current as the documents, and organisations refresh their content far less often than they assume. That case needs no retraining to explain and no retraining to fix.
Does this apply to assistants as well as to trained models?
Yes, and it is arguably more relevant there, because the material behind an assistant changes constantly while the model behind it changes on somebody else's schedule. Both move, neither is under your control, and only the combination is visible in what a user experiences.
07

Key takeaways

  • You changed nothing; the inputs, the sources, the world or the model underneath moved.
  • Stale sources are the case most buyers actually meet, and need no retraining to explain.
  • It is a movement, so detecting it requires a baseline recorded beforehand.
  • Gradual by nature, and users adapt to it faster than they report it.
09

Tools that use this

  • Glean

    Answer quality inheriting the estate's hygiene, stale documents included.

  • Notion AI

    Answers tracking whatever the workspace currently says.

  • Relevance AI

    Supervision preceding trust, which keeps a slow decline visible.

Last checked July 2026

All glossary terms