> Source: https://wexa.ai/docs/concepts/observability

# Observability

An **observability trace** is the end-to-end record of one request as it crosses the platform's
services. It is where you look to see which step took the time, and where a failure started rather
than where it surfaced.

Wexa emits traces, metrics and logs in OpenTelemetry's formats, over OTLP, to a collector you
provide. Nothing is stored by the platform itself, and nothing is served back over the Wexa
API — the trace goes to your collector, and you read it there.

## What you get without configuring anything

Two things are always true, whether or not telemetry export is switched on.

**Trace context propagates.** The gateway installs the W3C trace-context propagator unconditionally.
An inbound `traceparent` is honoured, and the gateway's own outbound calls to the services behind it
carry the context onward. A deployment where only some services export still produces connected
traces.

**Every response names its trace.** The gateway sets `X-Trace-Id` and a `Server-Timing` entry on the
response before the handler runs, so the identifier is present on successes, on refusals and on
requests that never authenticated. It is the value to quote when reporting a problem, and the TypeScript
SDK attaches it to the error object it throws.

## The shape of a trace

One tool call produces a nest of spans, in three layers.

**`gateway.request`** is the root: one server span per HTTP request. It is opened by middleware
installed *before* authentication, so a rejected or unauthenticated request is traced too — which is
what makes a trace useful for diagnosing "my key does not work". It carries the method and path, the
client address as resolved through any proxy headers, the origin and user agent, which kind of
credential was presented, whether the request arrived over the Wexa MCP server or the REST API, and
the response status.

**`lifecycle.<tool>`** is the tool call itself, a child of the request span. It carries the
`lifecycle_id`, the tool name, the surface, the project mode, the organization, project, actor and
role, the originating client and the connection the call arrived on, a redacted rendering of the
arguments, and — when the call was stopped — the reason and the stage that stopped it, **on the parent
span**. You do not have to hunt through children to learn why a call failed.

**`lifecycle.S1:authenticate` … `lifecycle.S10:record`** are one child span per stage of the
[lifecycle](/docs/concepts/lifecycle-and-audit), so the whole governed path renders as a waterfall
rather than a flat list. Each carries its stage name, its status, a sequence number, and stage-specific
detail: the quota counts and ceilings the call saw at `S3`, the rule, verdict, reasoning and decision
identifier at `S5`, the approval identifier and reason at `S6`. Each also carries a redacted input and
output for that stage, so you can see what entered and what left each checkpoint.

Two details of the shape are deliberate and worth knowing.

The `S8:execute` span is opened live, while the work happens, rather than replayed at the end. That is
what makes the downstream service call nest *inside* `S8` instead of appearing beside it — so the time
a tool spent waiting on the graph is visibly inside the execute stage.

Governance stages run in well under a millisecond, so their start timestamps collapse onto the same
instant. The sequence number on each stage span is what orders them; a viewer that sorts siblings by
timestamp alone will show them jumbled.

## How far a trace reaches

A trace covers the path from the caller, through the gateway, and into whatever does the work — the
context graph, the knowledge base, an execution. Inside the gateway you get a span for the request,
a span for the lifecycle and a span for each of its stages.

Not everything Wexa calls is traced. Where it is not, the hop still appears, as the gateway's own
client span with nothing nested under it. You can see that the call happened and how long it took;
you cannot see what happened inside it. Resolving a role while checking a grant is the common
example.

One thing reads traces rather than serving calls, so it never appears inside one: the analytics that
the console's observability panels are built from.

## Logs and metrics

**Logs are trace-correlated.** The gateway emits a structured log at each lifecycle stage carrying the
`lifecycle_id`, the stage, its status and detail, the tool and the organization, project and actor, plus
an ingress and a completion log for each request. Because they are emitted against the active span's
context, each record is stamped with the trace identifier, so a log line and its trace join up without
your having to correlate them by time.

**Two request metrics are recorded:** a counter of gateway requests and a histogram of request duration,
both broken down by route and response status. They come from the same middleware as the root span, so
they cover refused requests as well as served ones.

## Using a trace to debug a run

### A slow run

Open the trace by its identifier and read down the waterfall.

1. **Compare the root span's duration with `lifecycle.<tool>`.** A large gap means the time went
   somewhere other than the tool call — usually authentication or a role lookup, which
   appears as a client span with nothing beneath it.
2. **Look at which stage is wide.** The governance stages `S1`–`S7` and `S9`–`S10` are sub-millisecond
   by nature. If one of them is wide, that is the finding, not a rounding artefact.
3. **If `S8:execute` is the wide one, descend into it.** The downstream call nests there — the graph
   read behind [`query-context`](/docs/tools/query-context), the agent run behind
   [`run-agent`](/docs/tools/run-agent). A wide `S8` with a wide child is a slow service; a wide `S8`
   with a narrow child is time spent in the gateway around the call.
4. **Check the quota counts on `S3`.** A project running near its ceiling shows there, and it explains
   a run that is slow because its callers are backing off rather than because anything is.
5. **For a multi-call run, sort by the `lifecycle_id`, not by the tool.** An agent making twenty calls
   produces twenty traces; the one you want is the one whose identifier the caller reported.

### A failing run

The failure reason is on the parent span, so start there rather than at the deepest red child.

- A stage marked failed with `S5:policy` means a rule refused it — the rule name and its reasoning are
  on that stage's span, and the same decision is retrievable at `GET /v1/policy-decisions`.
- `S3:rate-quota` means a [ceiling](/docs/concepts/quota-and-credits) was hit; the counts on the stage
  say which bucket.
- `S6:approval` on a call that has a resume token means the token was rejected — spent, expired, or
  bound to a different tool, project or user.
- A pending status at `S6` is not a failure. The call parked at an
  [approval](/docs/concepts/policy-and-approvals) and the trace ends there legitimately; the rest of the
  story is in the trace of the resumed call.
- `S8:execute` failed means the tool itself or the service behind it failed. The child span carries the
  error, and its own status tells you whether the gateway gave up waiting or the service answered badly.

### When there is no trace

If the `X-Trace-Id` you were given finds nothing in your collector, the likely causes, in order: export
is off for that deployment; the collector endpoint is not reachable from the gateway, which degrades
silently by design so telemetry can never take the platform down; or the trace was not sampled. Head
sampling is configurable and samples everything by default, so a deployment that has lowered it will be
missing traces in proportion.

## What is redacted before it leaves

Arguments and results reaching a span are redacted first. Argument values whose key looks like a
credential — anything containing token, secret, password, key, auth, credential, cookie or session — are
masked, and result previews are truncated. The data-flow figures a span carries are structural: how many
records came in, how many went out, and the top-level field *names* of the result. Never the values.

Where a call touched the context graph, the nodes and relationships are recorded by identity only — the
label, the name of the identity key, and that key's reference value — so the trace can show which
entities a call moved without carrying what was in them.

## Where to look next

**Governance — Lifecycle and audit**
The ten stages the waterfall renders, and the durable account of the same call.

**Governance — Policy decisions and approvals**
What a refusal at `S5` means, and how to read the decision behind it.

**Governance — Quota and credits**
Reading the counts recorded at `S3`, and what a `429` in a trace is telling you.

**Knowledge and data — Where your data goes**
Which store an execute stage was talking to.
