Glossary
Knowledge base
A knowledge base is the defined set of material an assistant is allowed to answer from, chosen deliberately rather than gathered by default: policies, manuals, past tickets, whatever you decided to put in.
In plain terms
It is the shelf you point the assistant at. Everything on the shelf is fair game for an answer and everything off it is invisible, which sounds simple until you try to decide what goes on. That decision is editorial rather than technical, and in many organisations it is the first occasion on which somebody has had to say out loud which version of the policy is the current one.
Why it matters
An assistant answering from your documents is exactly as good as your documents, and this is where most disappointments actually originate. The common experience is not that the model is stupid; it is that three versions of the refund policy exist in three places, the assistant found all of them, and answered from whichever happened to rank highest. That is your filing being described accurately for the first time, and it is fixable in a way a model limitation is not.
How it works
The defining act is deciding what is in it. Which documents, which versions, whose word counts as final, and what happens to the things nobody wants to delete. None of that is a technical question, and skipping it is why so many of these projects stall at the point where somebody has to choose between two policies with the same title.
Contradiction is the biggest practical problem and the least discussed. Where two documents disagree, retrieval surfaces whichever sits closest to the question, and the answer is written confidently with no indication that a conflicting statement existed. You will not be told there was a disagreement. You will simply get one side of it.
Currency has to be managed deliberately, because nothing expires by itself. A superseded document that was never deleted goes on answering questions indefinitely, and the older it is the more likely it is to be phrased with the confidence of something that was once true. An owner and a review date do more for answer quality here than any change of tooling.
Structure helps considerably more than people expect. Clear headings, a date, sections short enough to stand alone, and an explicit line saying what a document replaces all make the difference between a passage that retrieves usefully and one that arrives without the context that made it meaningful. Writing for retrieval is a small discipline with a large effect.
Scope and permission decide who may ask what. One base serving everybody is simple and means every answer is drawn from material everybody may see, which is usually too restrictive. Separating by department is generally easier to run than per-document rules, and either way the checks belong at the point material is retrieved rather than in the wording of the request.
One question about returns, four things it can match
Seen in the wild
Load a deliberately chosen handful of documents into a notebook tool and notice how differently it behaves from a general assistant on the same questions.
NotebookLMAsk a question of a notes workspace and get an answer from pages your own team wrote, which is a knowledge base you have been building for years without calling it one.
Notion AISearch across a whole company's systems and see what happens when the same rule is stated in four places by four teams.
Glean
Common misconceptions
People assume
We will just point it at the shared drive.
In fact
That is the quickest route to confident answers drawn from a draft nobody has opened in years. A shared drive is an accumulation rather than a selection, and everything in it competes on equal terms. Curation is not a preliminary to the project; it substantially is the project.
People assume
More documents make it better.
In fact
More consistent documents make it better. More contradictory ones make it worse, because every addition is another candidate answer with no way of knowing which supersedes which. Removing three stale files usually improves results more than adding thirty good ones.
People assume
This is a technical project.
In fact
The technical part is largely solved and takes days. The part that takes months is deciding what is true, who owns each document, what happens to the old versions and who arbitrates when two departments disagree. Staffing it as an engineering job is the reliable way to be surprised.
Telling them apart
Knowledge base vs Training data
Knowledge base
What you gave it to answer from. Chosen by you, editable this afternoon, and checkable against any answer it produces.
What the model learnt from before release. Fixed, not published in detail, and nothing to do with you.
An answer that is wrong about the world came from the second. An answer that is wrong about your business came from the first, and only the first has an owner you can go and talk to.
Knowledge base vs RAG
Knowledge base
The material: which documents, which versions, kept current by whom.
The technique: searching that material at the moment a question is asked and supplying what it finds.
Vendors demonstrate the second and you live with the first. When a system disappoints after a good demonstration, look at the shelf before you look at the retrieval.
Questions
- What should go in it?
- The material somebody would reach for to answer the questions you actually get, in its current version only. Start narrow, with one subject and its genuinely authoritative documents, and expand once that works. Everything you add is another candidate answer, which is a reason to be reluctant rather than generous.
- Our documents contradict each other. What happens?
- You get one of them, chosen by whichever sat closest to the question, presented without any hint that the other exists. This is a frequent cause of a confidently wrong answer from an otherwise well-built system, and the fix is editorial: decide which is current and remove or clearly mark the other.
- Who should own it?
- The team whose work the documents describe, not whoever built the assistant. Ownership means deciding what is current, reviewing on a schedule and arbitrating when two versions disagree. An unowned base degrades quietly, and the first sign is usually a customer being told something that stopped being true a year ago.
- How do we keep it current?
- Give every document an owner and a review date, and make deleting the superseded version part of publishing the new one. That is unglamorous and it is the whole answer. No amount of retrieval sophistication compensates for a shelf where the old edition is still sitting next to the new one.
Key takeaways
- It is the defined set of material an assistant may answer from, and defining it is the actual work.
- Contradictory documents produce a confident answer from one side with no sign the other existed.
- Nothing expires on its own, so superseded material keeps answering until somebody removes it.
- Writing for retrieval, with dates, headings and self-contained sections, has a large effect.
- This is an editorial project with a technical component, and staffing it the other way round goes badly.
Tools that use this
- NotebookLM
A small, deliberately chosen shelf, which is the idea at a size you can hold.
- Notion AI
A knowledge base most teams have been building without naming it one.
- Glean
Company-wide, where the same rule stated four times becomes visible.
Last checked July 2026