> Source: https://wexa.ai/docs/concepts/where-your-data-goes

# Where your data goes

A project holds four stores. They are the single most common source of confusion for people new to
Wexa, because all four can loosely be described as "where the information is". They are not
interchangeable, and choosing the wrong one is the mistake that costs the most time to discover:
nothing fails at the moment you make it, and the cost arrives later as a question none of your
stores can answer.

Each of the four has exactly one property that no other has. That is the fastest way to tell them
apart, and it is how the rest of this page is organized.

  <Card title="Context graph" eyebrow="What is connected to what" href="/docs/concepts/context-graph">
    Entities and their relationships. The only store you query by traversing.
  </Card>
  <Card title="Knowledge base" eyebrow="What did we write about this" href="/docs/concepts/knowledge-base">
    Documents and passages, retrieved by meaning. The only store holding prose.
  </Card>
  <Card title="Data catalog" eyebrow="Where did this table come from" href="/docs/concepts/data-catalog">
    Your data assets and their lineage, quality and ownership — described, not copied.
  </Card>
  <Card title="Project ontology" eyebrow="What kinds of thing may exist" href="/docs/concepts/project-ontology">
    Your project's vocabulary of node and relationship types. Shape, not content.
  </Card>

## The context graph

**What it stores.** The nodes and relationships describing your project's world — its entities and
how they relate. A customer, an order, a repository, a person, and the edges between them.

**What it is good at.** Questions whose answer depends on connection. Which customer this order
belongs to, which services depend on the one you are about to change, who touched a record and in
what order. It is the only one of the four that is a graph, and the only one you query by
traversing relationships, so it is the only store that can answer a question of that shape at all.

**What it is the wrong choice for.** Long prose that you want answered approximately. You can put a
paragraph on a node, but nothing will search it by meaning, and a question phrased differently from
the text will not find it. It is also the wrong place to describe a table that lives in your
warehouse: you would be modelling a copy of something the data catalog already tracks properly,
which then quietly goes stale.

## The knowledge base

**What it stores.** Documents and passages, held for retrieval by meaning rather than by exact
match. It is the only store holding unstructured prose.

**What it is good at.** Getting a grounded answer out of writing nobody wants to read in full —
policies, runbooks, specifications, support transcripts, past decisions. You ask a question in your
own words and get back the passages that are about it, whether or not they use your wording.

**What it is the wrong choice for.** Anything where approximate is not good enough. A retrieved
passage is evidence, not a lookup: if you need the current credit limit for one customer, the
answer belongs in the context graph, where there is exactly one of it. It is also the wrong place
for facts you will want to traverse — a passage saying two systems are connected cannot be walked
from one to the other.

## The data catalog

**What it stores.** The inventory of your data assets — tables, columns, dashboards — together
with their lineage, their quality scorecards and their ownership. It is the only store describing
data that lives *outside* Wexa, and it holds a description of that data rather than the data
itself.

**What it is good at.** Deciding whether to trust something before you use it. Where a table came
from, what fed it, who owns it, what its quality judgement says. That is the question that arrives
just before someone builds a report on a number they have never questioned.

**What it is the wrong choice for.** The rows. Cataloguing an asset does not copy its contents into
Wexa, so "put the table in the catalog" will not make the values queryable — you will have a
faithful description of data you still cannot read. It is also the wrong place for entities that
were born inside Wexa and live nowhere else; those are context.

## The project ontology

**What it stores.** Your project's vocabulary of node and relationship types, layered over the
platform schema that every project shares.

**What it is good at.** Saying what *kinds* of thing may exist. It is the only one of the four that
describes the shape of another store rather than holding content of its own: it governs what can
appear in the context graph.

**What it is the wrong choice for.** Anything with an instance. There is no such thing as putting
a customer in the project ontology — you put the *type* `Customer` there, and the customer in the
context graph. If you find yourself adding a type for a single real thing, the type belongs to the
vocabulary and the thing belongs in the context graph.

## Where should I put this?

Four concrete cases, worked through.

### A 40-page procurement policy that people keep asking questions about

**The knowledge base.** It is prose, and the questions will be phrased in a hundred ways the
document never was. Put it in the context graph and you get a long string on a node nobody can
search; put it in the data catalog and you have described a document you still cannot read.

### "Which customer does this order belong to, and which agent last touched it?"

**The context graph.** The question is entirely about connection, and traversal is the only way to
answer it. Two hops, two relationships. No amount of similarity search produces this answer,
because the answer is not written down anywhere as a sentence.

### A warehouse table your team is about to build a report on

**The data catalog.** You want its lineage, its owner and its quality scorecard — the things that
tell you whether the number is safe to publish. The table itself stays where it is; the catalog
describes it. If you also want your agents to reason about that table as an entity in your world,
that relationship is context, and the two records coexist.

### "We now track Contracts, and a Contract renews a Subscription — neither type exists yet"

**The project ontology.** You are describing what kinds of thing may exist, which is the one thing
the other three stores do not do. Add the types first; the individual contracts then go into the
context graph as nodes of that type.

## When the answer is more than one

Some material legitimately lands in two places, and that is not a sign you have chosen wrongly. A
support transcript is prose — it belongs in the knowledge base — while the customer, the product
and the escalation it mentions are entities, and they belong in the context graph. What you should
not do is copy the same thing into both and expect them to stay in step: store each part where its
questions will be asked.

## Why they are separate at all

Because they are governed the same way but queried differently. Keeping them separate means a query
does one thing well instead of four things approximately, while policy decisions, approvals, audit
and quota apply identically across all four. You get the difference where it helps, and none of it
where it does not.
