On this page

Observability

An observability trace is the end-to-end record of one request as it crosses the platform's services. It is where you look to see which step took the time, and where a failure started rather than where it surfaced.

Wexa emits traces, metrics and logs in OpenTelemetry's formats, over OTLP, to a collector you provide. Nothing is stored by the platform itself, and nothing is served back over the Wexa API — the trace goes to your collector, and you read it there.

What you get without configuring anything

Two things are always true, whether or not telemetry export is switched on.

Trace context propagates. The gateway installs the W3C trace-context propagator unconditionally. An inbound traceparent is honoured, and the gateway's own outbound calls to the services behind it carry the context onward. A deployment where only some services export still produces connected traces.

Every response names its trace. The gateway sets X-Trace-Id and a Server-Timing entry on the response before the handler runs, so the identifier is present on successes, on refusals and on requests that never authenticated. It is the value to quote when reporting a problem, and the TypeScript SDK attaches it to the error object it throws.

The shape of a trace

One tool call produces a nest of spans, in three layers.

gateway.request is the root: one server span per HTTP request. It is opened by middleware installed before authentication, so a rejected or unauthenticated request is traced too — which is what makes a trace useful for diagnosing "my key does not work". It carries the method and path, the client address as resolved through any proxy headers, the origin and user agent, which kind of credential was presented, whether the request arrived over the Wexa MCP server or the REST API, and the response status.

lifecycle.<tool> is the tool call itself, a child of the request span. It carries the lifecycle_id, the tool name, the surface, the project mode, the organization, project, actor and role, the originating client and the connection the call arrived on, a redacted rendering of the arguments, and — when the call was stopped — the reason and the stage that stopped it, on the parent span. You do not have to hunt through children to learn why a call failed.

lifecycle.S1:authenticate … lifecycle.S10:record are one child span per stage of the lifecycle, so the whole governed path renders as a waterfall rather than a flat list. Each carries its stage name, its status, a sequence number, and stage-specific detail: the quota counts and ceilings the call saw at S3, the rule, verdict, reasoning and decision identifier at S5, the approval identifier and reason at S6. Each also carries a redacted input and output for that stage, so you can see what entered and what left each checkpoint.

Two details of the shape are deliberate and worth knowing.

The S8:execute span is opened live, while the work happens, rather than replayed at the end. That is what makes the downstream service call nest inside S8 instead of appearing beside it — so the time a tool spent waiting on the graph is visibly inside the execute stage.

Governance stages run in well under a millisecond, so their start timestamps collapse onto the same instant. The sequence number on each stage span is what orders them; a viewer that sorts siblings by timestamp alone will show them jumbled.

How far a trace reaches

A trace covers the path from the caller, through the gateway, and into whatever does the work — the context graph, the knowledge base, an execution. Inside the gateway you get a span for the request, a span for the lifecycle and a span for each of its stages.

Not everything Wexa calls is traced. Where it is not, the hop still appears, as the gateway's own client span with nothing nested under it. You can see that the call happened and how long it took; you cannot see what happened inside it. Resolving a role while checking a grant is the common example.

One thing reads traces rather than serving calls, so it never appears inside one: the analytics that the console's observability panels are built from.

Logs and metrics

Logs are trace-correlated. The gateway emits a structured log at each lifecycle stage carrying the lifecycle_id, the stage, its status and detail, the tool and the organization, project and actor, plus an ingress and a completion log for each request. Because they are emitted against the active span's context, each record is stamped with the trace identifier, so a log line and its trace join up without your having to correlate them by time.

Two request metrics are recorded: a counter of gateway requests and a histogram of request duration, both broken down by route and response status. They come from the same middleware as the root span, so they cover refused requests as well as served ones.

Using a trace to debug a run

A slow run

Open the trace by its identifier and read down the waterfall.

  1. Compare the root span's duration with lifecycle.<tool>. A large gap means the time went somewhere other than the tool call — usually authentication or a role lookup, which appears as a client span with nothing beneath it.
  2. Look at which stage is wide. The governance stages S1–S7 and S9–S10 are sub-millisecond by nature. If one of them is wide, that is the finding, not a rounding artefact.
  3. If S8:execute is the wide one, descend into it. The downstream call nests there — the graph read behind query-context, the agent run behind run-agent. A wide S8 with a wide child is a slow service; a wide S8 with a narrow child is time spent in the gateway around the call.
  4. Check the quota counts on S3. A project running near its ceiling shows there, and it explains a run that is slow because its callers are backing off rather than because anything is.
  5. For a multi-call run, sort by the lifecycle_id, not by the tool. An agent making twenty calls produces twenty traces; the one you want is the one whose identifier the caller reported.

A failing run

The failure reason is on the parent span, so start there rather than at the deepest red child.

  • A stage marked failed with S5:policy means a rule refused it — the rule name and its reasoning are on that stage's span, and the same decision is retrievable at GET /v1/policy-decisions.
  • S3:rate-quota means a ceiling was hit; the counts on the stage say which bucket.
  • S6:approval on a call that has a resume token means the token was rejected — spent, expired, or bound to a different tool, project or user.
  • A pending status at S6 is not a failure. The call parked at an approval and the trace ends there legitimately; the rest of the story is in the trace of the resumed call.
  • S8:execute failed means the tool itself or the service behind it failed. The child span carries the error, and its own status tells you whether the gateway gave up waiting or the service answered badly.

When there is no trace

If the X-Trace-Id you were given finds nothing in your collector, the likely causes, in order: export is off for that deployment; the collector endpoint is not reachable from the gateway, which degrades silently by design so telemetry can never take the platform down; or the trace was not sampled. Head sampling is configurable and samples everything by default, so a deployment that has lowered it will be missing traces in proportion.

What is redacted before it leaves

Arguments and results reaching a span are redacted first. Argument values whose key looks like a credential — anything containing token, secret, password, key, auth, credential, cookie or session — are masked, and result previews are truncated. The data-flow figures a span carries are structural: how many records came in, how many went out, and the top-level field names of the result. Never the values.

Where a call touched the context graph, the nodes and relationships are recorded by identity only — the label, the name of the identity key, and that key's reference value — so the trace can show which entities a call moved without carrying what was in them.

Where to look next