Data catalog
The data catalog is the inventory of a customer's data assets — tables, columns, dashboards — together with their lineage, their quality and their ownership. It is the only one of a project's four stores that describes data living outside Wexa, and it holds a description of that data rather than the data itself. If you have not yet chosen between the four, start at where your data goes.
Assets
One entry in the catalog is a catalog asset. The platform models a handful of kinds, and they
nest the way your warehouse does: a DataService is a source such as a warehouse, a database or a
stream; a Database is a logical database inside it; a Table is a relation, a view or a streaming
topic; a Column is a field within a table. Alongside those sit Dashboard for a report or chart,
Pipeline for a job that moves or transforms data, MLModel for a trained model, GlossaryTerm
for a business definition, Classification for a tag such as PII, and DataOwner for the person
answerable for an asset.
An asset records what the thing is, where it lives and who is answerable for it. It does not record the rows. Asking the catalog for the contents of a table is asking the wrong store — the catalog can tell you where that table is and whether you should trust it, and connector-read is what fetches the rows.
Lineage
Lineage is the recorded path data took from one catalog asset to another. It is recorded at two grains: table to table, and column to column. The finer grain is the one that answers the question people actually ask, because "this report is wrong" is usually about one column rather than a whole table.
Read one way it answers "where did this come from". Read the other way it answers "what breaks if I
change this". Asset-producing jobs are modelled too: a Pipeline or an MLModel that writes a table
or a dashboard is recorded as feeding it, so a job shows up as a step in the chain rather than as
an unexplained gap in it.
Lineage is a property of the catalog rather than of the context graph: it describes movement between assets outside Wexa, not relationships between entities inside it.
Scorecards
A scorecard is a quality judgement attached to a catalog asset. It is what turns the catalog from a list of things that exist into something you can act on, because it carries whether an asset is fit to use as well as whether it is there.
The judgements are computed from the catalog itself rather than entered by hand. They are the questions a data operator would otherwise ask one asset at a time:
- columns classified as PII that carry no description — the one graded as a warning, because an undocumented sensitive field is a decision nobody can review;
- columns with no parent table, which usually means an incomplete sync rather than a real orphan;
- tables with no lineage recorded at all, so nothing is known about where they came from;
- tables with no description;
- tables with no recorded owner.
Each rule comes back with how many assets fail it and a sample of which ones, and the report carries an overall status so you can tell at a glance whether anything needs attention.
What backs it
The catalog's contents are produced by OpenMetadata, which the platform runs as the engine that crawls your sources and computes lineage. Its output does not stay there: it is translated into catalog assets and relationships and stored in Wexa, under the same project isolation, the same policy decisions and the same audit record as everything else. That is why the catalog is queryable next to your context rather than in a separate tool, and it is the only place in this documentation where that system is named — everywhere else, the concept is the data catalog.
Two consequences worth knowing. Synchronization is periodic, with a configurable cadence, and a deployment can also be sent an update when a source changes rather than waiting for the next pass; there is a call to trigger a pass on demand and a call to read the statistics from the last one. And classifications flow onwards: a tag such as PII arriving from a source updates the restricted set that Wexa's policy engine enforces, so classifying a column at the source tightens what agents may do with it here.
Reading the catalog
The catalog is read over the REST API rather than through a tool. There are calls to search assets, to read one asset in detail, to read an asset's lineage, to read the catalog's own type definitions, and to read the scorecard report for a project.
Assets also appear in the context graph as nodes, which means
query-context can traverse them — including to find which of them are
live-readable, which is the first step of a connector read.
How it is filled
Through ingestion: a source is connected by provisioning a connector, and ingestion brings its metadata in so it appears in the catalog and its entities appear in the context graph.
Source code is a separate path with separate concepts. A repository connection links a project to a source-code repository, and code sync keeps the code graph in step with it. A repository connection is not a connector and is never provisioned.