> Source: https://wexa.ai/docs/concepts/knowledge-base

# Knowledge base

The **knowledge base** holds documents and passages stored for retrieval by meaning rather than by
exact match. It is the only one of a project's four stores that holds unstructured prose, and the
only one where a question phrased in your own words can find writing that never used those words.
If you have not yet chosen between the four, start at
[where your data goes](/docs/concepts/where-your-data-goes).

Every deployment lists the knowledge-base tools. Where the service behind them is not configured,
[`knowledge-base-add`](/docs/tools/knowledge-base-add) and
[`knowledge-base-retrieve`](/docs/tools/knowledge-base-retrieve) are present and failing, with
`is not configured (no data-service URL)`, rather than absent from the tool list.

## There is no knowledge base to create

This is the first thing that surprises people, and it saves a search for a creation call that does
not exist. A knowledge base is not an object you provision. It is the project's stored content
filtered by **tag**. Adding a document under a tag nobody has used before is all it takes to start
one, and there is no corresponding step to delete a knowledge base — there are only documents and
the tags they carry.

That is why tags are required on both sides of the operation, and why they are worth choosing
carefully rather than treating as decoration.

## Adding a document

[`knowledge-base-add`](/docs/tools/knowledge-base-add) takes the document text and a non-empty list
of tags. Both are required, and an add with no tags is refused rather than accepted quietly, because
a document stored with no tags could never be found again — retrieval has no way to reach it.

The project the document lands in comes from your scope, not from anything you send.

## Retrieving

[`knowledge-base-retrieve`](/docs/tools/knowledge-base-retrieve) works in two stages, and
understanding the order is the difference between getting results and getting an empty list.

**First, tags narrow the corpus.** Tags are required. Retrieval searches the documents filed under
the tags you pass and nothing else. Passing a tag that nothing was stored under returns nothing, no
matter how well-phrased the question is.

**Then, meaning ranks what is left.** An optional `goal` — a question in ordinary language — reorders
the narrowed set so the passages that are *about* your question come first. The goal does not widen
the search; it only decides the order within the tags you already named. Omit it and you get the
tagged documents without that ranking.

The rest is mechanical: a result limit, a cursor for paging through more than one page of results,
and a switch for whether image payloads come back inline.

## How agents use it

An agent reads this corpus when it has been configured to, and the tags it may read are part of
that configuration. So the same tag vocabulary that scopes your own retrieval also scopes what an
agent can see: giving an agent narrower tags is how you give it a narrower corpus, without moving
any documents. See [agents](/docs/concepts/agents).

## How it differs from the context graph

Both stores can hold something about the same subject, and the distinction is not about the
subject — it is about the shape of the question.

| | Knowledge base | [Context graph](/docs/concepts/context-graph) |
|---|---|---|
| Holds | Prose, as written | Entities and their relationships |
| Reached by | Tags, then ranked by meaning | Traversing relationships in a query |
| Answers | "What did we write about this?" | "What is this connected to?" |
| Returns | Passages, as evidence | Rows, as facts |
| Wrong for | A single current value | Long prose answered approximately |

A retrieved passage is **evidence**, not a lookup. Similarity ranking always returns its closest
match, which means it returns something even when nothing is really relevant. That is a feature when
the corpus is prose and a hazard when the question has exactly one correct answer. If you need the
current credit limit for one customer, that belongs in the context graph, where there is exactly one
of it and the answer is either there or absent.

The reverse mistake is subtler. A passage saying that two systems are connected cannot be walked
from one to the other. Facts you will want to traverse belong in the graph even when they arrived
inside a document — which is why a support transcript often lands in both stores: the prose here,
and the customer, product and escalation it names as entities over there.

## Governance

Adding and retrieving are governed like every other call. Each passes a policy decision first,
draws on the project's quota, and is written to the audit record. That last part matters more here
than elsewhere: "what did this agent read" is a question the audit record answers directly, rather
than one you have to reconstruct from what the agent then said. See
[policy and approvals](/docs/concepts/policy-and-approvals).
