On this page

Ingestion

Ingestion is bringing a connected source's data into Wexa so that it appears in the context graph and the data catalog. Provisioning a connector makes a source reachable; ingestion is what makes it present. They are separate steps, and a project that has done the first and not the second has a connector and an empty graph.

What ingestion produces

One pass over a source produces two different kinds of record, which is why ingestion touches two stores rather than one.

Metadata becomes catalog assets. The structure of the source — its databases, tables, columns, dashboards and models, together with the lineage between them and any classifications and glossary terms attached to them — lands as catalog assets and the relationships between them.

Content becomes context. The entities a source is actually about — the messages, tickets, records, people — land as nodes and relationships in the context graph, typed according to the project's vocabulary.

Both land under the calling scope's project. Neither is a copy you then have to keep in step by hand: ingestion is repeatable, and repeated passes merge onto the same nodes rather than accumulating duplicates.

Why it is asynchronous

Ingestion crosses a network to a system Wexa does not control, and how long it takes is a property of that system rather than of Wexa. A source with a hundred tables and a source with a hundred thousand differ by orders of magnitude, and holding an HTTP request open for the second one would mean a caller whose only signal is a timeout.

So an ingestion call answers as soon as the work is accepted, and the work continues after the response has been sent. That shifts a burden onto the caller — you now have to ask whether it finished — and the two paths below each give you a way to do that.

Checking a catalog synchronization

A catalog synchronization is triggered on demand and reports back separately. The trigger call requires a catalog-write grant and returns immediately, leaving the pass running behind it. Reading the synchronization statistics afterwards is what tells you what happened: how many tables and columns were seen, how many assets were created and how many updated, how many lineage edges were recorded at each grain, and how many edges were rejected.

Three answers to the trigger are worth recognizing, because each means something different:

  • Accepted. The pass is running. Read the statistics to see the result.
  • Already running. A pass is in flight, and a second one is refused rather than started alongside it. The response says to watch the statistics for the result of the pass that is already going.
  • Not configured. No synchronization exists for this organization, so there is nothing to run. This is the answer that used to be the most expensive to misread — it is not a slow sync, it is no sync, and it needs an administrator rather than patience.

A deployment can also be notified when a source changes, rather than waiting for the next scheduled pass, and the cadence of those scheduled passes is a deployment setting.

Checking a code-sync ingestion job

The code sync path has an explicit job model, and it is the clearest example of the shape. A bulk ingestion is submitted, the response carries a job identifier, and a separate call reads that job's status by identifier. Poll it until the job reports that it is finished.

Everything about the request is pinned to your scope on the way through: the organization and project the data lands in are taken from your token and overwritten on the payload, so a client cannot ingest into a project it was not granted, whatever it sends.

Ingestion is not the only way to read a source

It is worth knowing when not to ingest.

connector-read runs a read-only query against the source live, and copies nothing into Wexa. For a large warehouse or database, that is usually the better answer: you get current rows without maintaining a second copy of them, and without paying to ingest data you will read once.

The division that works in practice is: ingest the structure, read the rows live. The catalog gives you the shape of the source, which is small, slow-changing and worth having locally so agents can reason about it without a round trip. The rows are large, fast-changing and best fetched when they are actually wanted. See connectors.

Governance

Ingestion is governed like everything else. A pass runs under a scope, requires the relevant grant, draws on the project's quota and writes to the audit record. Classifications arriving from a source feed onwards into the restricted set that the policy engine enforces, so labelling a column as sensitive at the source tightens what agents may do with it in Wexa. See policy and approvals.