Observability
An observability trace is the end-to-end record of one request as it crosses the platform's services. It is where you look to see which step took the time, and where a failure started rather than where it surfaced.
Wexa emits traces, metrics and logs in OpenTelemetry's formats, over OTLP, to a collector you provide. Nothing is stored by the platform itself, and nothing is served back over the Wexa API — the trace goes to your collector, and you read it there.
What you get without configuring anything
Two things are always true, whether or not telemetry export is switched on.
Trace context propagates. The gateway installs the W3C trace-context propagator unconditionally.
An inbound traceparent is honoured, and the gateway's own outbound calls to the services behind it
carry the context onward. A deployment where only some services export still produces connected
traces.
Every response names its trace. The gateway sets X-Trace-Id and a Server-Timing entry on the
response before the handler runs, so the identifier is present on successes, on refusals and on
requests that never authenticated. It is the value to quote when reporting a problem, and the TypeScript
SDK attaches it to the error object it throws.
The shape of a trace
One tool call produces a nest of spans, in three layers.
gateway.request is the root: one server span per HTTP request. It is opened by middleware
installed before authentication, so a rejected or unauthenticated request is traced too — which is
what makes a trace useful for diagnosing "my key does not work". It carries the method and path, the
client address as resolved through any proxy headers, the origin and user agent, which kind of
credential was presented, whether the request arrived over the Wexa MCP server or the REST API, and
the response status.
lifecycle.<tool> is the tool call itself, a child of the request span. It carries the
lifecycle_id, the tool name, the surface, the project mode, the organization, project, actor and
role, the originating client and the connection the call arrived on, a redacted rendering of the
arguments, and — when the call was stopped — the reason and the stage that stopped it, on the parent
span. You do not have to hunt through children to learn why a call failed.
lifecycle.S1:authenticate … lifecycle.S10:record are one child span per stage of the
lifecycle, so the whole governed path renders as a waterfall
rather than a flat list. Each carries its stage name, its status, a sequence number, and stage-specific
detail: the quota counts and ceilings the call saw at S3, the rule, verdict, reasoning and decision
identifier at S5, the approval identifier and reason at S6. Each also carries a redacted input and
output for that stage, so you can see what entered and what left each checkpoint.
Two details of the shape are deliberate and worth knowing.
The S8:execute span is opened live, while the work happens, rather than replayed at the end. That is
what makes the downstream service call nest inside S8 instead of appearing beside it — so the time
a tool spent waiting on the graph is visibly inside the execute stage.
Governance stages run in well under a millisecond, so their start timestamps collapse onto the same instant. The sequence number on each stage span is what orders them; a viewer that sorts siblings by timestamp alone will show them jumbled.
How far a trace reaches
A trace covers the path from the caller, through the gateway, and into whatever does the work — the context graph, the knowledge base, an execution. Inside the gateway you get a span for the request, a span for the lifecycle and a span for each of its stages.
Not everything Wexa calls is traced. Where it is not, the hop still appears, as the gateway's own client span with nothing nested under it. You can see that the call happened and how long it took; you cannot see what happened inside it. Resolving a role while checking a grant is the common example.
One thing reads traces rather than serving calls, so it never appears inside one: the analytics that the console's observability panels are built from.
Logs and metrics
Logs are trace-correlated. The gateway emits a structured log at each lifecycle stage carrying the
lifecycle_id, the stage, its status and detail, the tool and the organization, project and actor, plus
an ingress and a completion log for each request. Because they are emitted against the active span's
context, each record is stamped with the trace identifier, so a log line and its trace join up without
your having to correlate them by time.
Two request metrics are recorded: a counter of gateway requests and a histogram of request duration, both broken down by route and response status. They come from the same middleware as the root span, so they cover refused requests as well as served ones.
Using a trace to debug a run
A slow run
Open the trace by its identifier and read down the waterfall.
- Compare the root span's duration with
lifecycle.<tool>. A large gap means the time went somewhere other than the tool call — usually authentication or a role lookup, which appears as a client span with nothing beneath it. - Look at which stage is wide. The governance stages
S1–S7andS9–S10are sub-millisecond by nature. If one of them is wide, that is the finding, not a rounding artefact. - If
S8:executeis the wide one, descend into it. The downstream call nests there — the graph read behindquery-context, the agent run behindrun-agent. A wideS8with a wide child is a slow service; a wideS8with a narrow child is time spent in the gateway around the call. - Check the quota counts on
S3. A project running near its ceiling shows there, and it explains a run that is slow because its callers are backing off rather than because anything is. - For a multi-call run, sort by the
lifecycle_id, not by the tool. An agent making twenty calls produces twenty traces; the one you want is the one whose identifier the caller reported.
A failing run
The failure reason is on the parent span, so start there rather than at the deepest red child.
- A stage marked failed with
S5:policymeans a rule refused it — the rule name and its reasoning are on that stage's span, and the same decision is retrievable atGET /v1/policy-decisions. S3:rate-quotameans a ceiling was hit; the counts on the stage say which bucket.S6:approvalon a call that has a resume token means the token was rejected — spent, expired, or bound to a different tool, project or user.- A pending status at
S6is not a failure. The call parked at an approval and the trace ends there legitimately; the rest of the story is in the trace of the resumed call. S8:executefailed means the tool itself or the service behind it failed. The child span carries the error, and its own status tells you whether the gateway gave up waiting or the service answered badly.
When there is no trace
If the X-Trace-Id you were given finds nothing in your collector, the likely causes, in order: export
is off for that deployment; the collector endpoint is not reachable from the gateway, which degrades
silently by design so telemetry can never take the platform down; or the trace was not sampled. Head
sampling is configurable and samples everything by default, so a deployment that has lowered it will be
missing traces in proportion.
What is redacted before it leaves
Arguments and results reaching a span are redacted first. Argument values whose key looks like a credential — anything containing token, secret, password, key, auth, credential, cookie or session — are masked, and result previews are truncated. The data-flow figures a span carries are structural: how many records came in, how many went out, and the top-level field names of the result. Never the values.
Where a call touched the context graph, the nodes and relationships are recorded by identity only — the label, the name of the identity key, and that key's reference value — so the trace can show which entities a call moved without carrying what was in them.
Where to look next
The ten stages the waterfall renders, and the durable account of the same call.
What a refusal at S5 means, and how to read the decision behind it.
Reading the counts recorded at S3, and what a 429 in a trace is telling you.
Which store an execute stage was talking to.