On this page

Error handling and retries

Both SDKs ship a complete error hierarchy with retryability already decided. If you do not know it is there you will write a retry loop that is worse than the one the library already offers — one that retries a validation failure forever, or gives up on the one class of failure that was always going to succeed on the second attempt.

This page is written for the TypeScript SDK, but the taxonomy, the retry policy and the resume boundary are identical in Python. Where the two differ, the difference is named.

The hierarchy

Fourteen classes in TypeScript, thirteen in Python. Every one of them extends WexaError, so a single catch can cover everything the library can throw.

WexaError                     base — carries stage, status, lifecycleId, traceId, raw, retryAfter
├── AuthError                 401  invalid_token, unauthorized
├── ForbiddenError            403  the credential lacks the grant, or S2 refused the scope
│   └── PolicyDenied          403  S5 — a policy rule refused the call outright
├── ValidationError           400/422  malformed body, rejected Cypher, failed dry run (S4/S7)
├── ApprovalError             403  S6 — the resume token is bad, expired, rejected or spent
├── QuotaExceeded             429  S3 — over the request window
├── NotFound                  404  no such approval, flow or execution in your scope
├── ConflictError             409  approval not pending, or agent not ready
├── UpstreamError             502  S8 — a downstream service refused or was unreachable
│   └── TimeoutError          504  an upstream deadline was exceeded
├── ConfigurationError        502/503  unconfigured, unavailable — a deployment fault
├── TransportError            nothing reached the gateway (TypeScript only)
└── ApprovalRequired          202  a person must approve before the call proceeds

The inheritance is load-bearing rather than cosmetic. Python's retry policy is isinstance(e, (UpstreamError, QuotaExceeded)), and the TypeScript port keeps exactly that membership, so which class extends which decides what retries. TimeoutError retries because it extends UpstreamError. ConfigurationError does not retry, despite carrying a 5xx status, because it extends WexaError directly.

ApprovalRequired and ApprovalError are siblings, not parent and child. The first means a person has not decided yet; the second means the token that would have resumed the call is no longer good.

What every error carries

FieldWhat it is
stageThe lifecycle stage that refused, S2:resolve-scope through S8:execute. undefined for refusals outside the lifecycle — a 401, a 404, unconfigured.
statusHTTP status, or undefined when nothing reached the gateway.
lifecycleIdThe governed lifecycle this call ran as. A 401 predates it, so it is absent there.
traceIdThe value of the response's X-Trace-Id header.
rawThe full decoded response body. On a TransportError it is the original thrown error instead.
retryAfterSeconds from a Retry-After header, clamped to [0, 60].
resumeSpentWhether a resume token was consumed before this refusal. Only meaningful on the resume path.

Quote lifecycleId and traceId in a support request. Both are appended to the error message as well, so a bare log line already carries them:

quota exceeded for project proj_c10 (retry after 35s) [lifecycle_id=qlc_000157 trace_id=36088b71f728852753f230188485b0dd]

That line is the SDK's real output, not an illustration.

How a response becomes a class

raiseFor maps a non-2xx response in three tiers, first match wins.

  1. The error string in the body. For a lifecycle refusal that string is the stage id, so S5:policy becomes PolicyDenied and S3:rate-quota becomes QuotaExceeded. The same table maps the non-lifecycle codes: invalid_token and unauthorized to AuthError, harness_unreachable and data_service_unreachable to UpstreamError, unconfigured and unavailable to ConfigurationError.
  2. The HTTP status. 400 and 422 to ValidationError, 401 to AuthError, 403 to ForbiddenError, 404 to NotFound, 409 to ConflictError, 429 to QuotaExceeded, 504 to TimeoutError.
  3. The fallback. Any 5xx becomes UpstreamError; anything else becomes a bare WexaError.

502 and 503 are deliberately missing from tier 2. That omission is the whole reason an unrecognised 502 is retryable while unconfigured is not: an unlabelled 502 falls through to the fallback and becomes a retryable UpstreamError, whereas a 503 that names itself unconfigured is caught in tier 1 and becomes a ConfigurationError that will never retry. Retrying a deployment fault does not fix it; it just costs you three attempts before you tell an operator.

ApprovalRequired appears in none of these tables. A 202 is built in the transport layer, which is the only place that holds the approval object and the tool name its constructor needs.

Catching them

Order the handlers from specific to general.

import {
  ApprovalRequired, PolicyDenied, ForbiddenError, ValidationError,
  QuotaExceeded, WexaError,
} from '@wexa-fabric/sdk'

try {
  await fabric.saveContext({ nodes })
} catch (e) {
  if (e instanceof ApprovalRequired) {
    await checkpoint(e.approvalId, e.resumeToken)   // durable, before you wait
  } else if (e instanceof PolicyDenied) {
    // S5 — a rule refused it. Terminal; a different credential will not help.
  } else if (e instanceof ForbiddenError) {
    // S2 — mint a credential that holds the grant.
  } else if (e instanceof ValidationError) {
    // Your payload. e.stage says whether the gateway or the SDK rejected it.
  } else if (e instanceof QuotaExceeded) {
    // S3 — wait e.retryAfter seconds, or let retry: true do it.
  } else if (WexaError.isWexaError(e)) {
    // Everything else the gateway sent, plus TransportError.
  } else {
    throw e
  }
}

WexaError.isWexaError(e) is a string-tag check rather than a prototype check. Prefer it to instanceof WexaError when two copies of the package could end up in one dependency tree: instanceof fails across duplicated installs and this does not. It has no Python equivalent, because Python cannot get two copies of a module into one interpreter the same way.

Which errors retry

Nothing retries unless you ask. The gateway has no idempotency key, so a retried runAgent starts a second real run — the library refuses to make that decision for you.

await fabric.queryContext({ query }, { retry: true })   // safe: it is a read

retries: 3 on the constructor means three attempts in total, not three extra ones. Only UpstreamError — and therefore TimeoutError — and QuotaExceeded are retried.

Every class in the hierarchy answers the two predicates like this:

ClassisRetryableisResumeRetryable
WexaErrorfalsefalse
AuthErrorfalsefalse
ForbiddenErrorfalsefalse
PolicyDeniedfalsefalse
ValidationErrorfalsefalse
ApprovalErrorfalsefalse
QuotaExceededtruetrue
NotFoundfalsefalse
ConflictErrorfalsefalse
UpstreamErrortruefalse
TimeoutErrortruefalse
ConfigurationErrorfalsefalse
TransportErrorfalsefalse
ApprovalRequiredfalsefalse

TransportError does not retry. Nothing reached the gateway, so the SDK cannot tell a wrong workspace URL from a network blip, and it declines to guess.

Backoff

When the gateway sends Retry-After, that number is honoured — it is computed from the quota window's real remaining time, and returning before the window resets only spends another 429. A small random offset is added even then, so several clients told to wait the same number of seconds do not all return in the same millisecond. With no header, backoff is full jitter — a uniform draw from [0, 2^attempt) — capped at the 60-second quota window.

Retry-After is parsed strictly. RFC 9110 permits an HTTP-date as well as a count of seconds; the gateway sends seconds, and the date form is ignored rather than misparsed into a wrong sleep. A non-integer is ignored for the same reason.

Here is that path executed end to end against a live gateway, after deliberately exhausting the project's 120-request window:

A no retry -> QuotaExceeded status 429 stage S3:rate-quota retryAfter 35 isRetryable true isResumeRetryable true
B retry:true succeeded after 35.9 s; lifecycleId qlc_000159

Plain retry is not resume retry

This is the distinction that makes the two predicates worth having, and it is the one most likely to cost you a write.

A call that gates for approval throws ApprovalRequired carrying a single-use resumeToken. When you send that token back, the gateway redeems it between S3 (quota) and S4 (validate).

S1 authenticate → S2 resolve-scope → S3 rate-quota → [REDEEM] → S4 validate-input → S5 policy → … → S8 execute
                                          ↑                          ↑
                          refused here: token intact      refused here: token gone

So a refusal at S3 arrives before the redemption and the token survives; every refusal after that point means the token has been burned, the write did not land, and re-running needs a fresh human approval. That is the entire content of isResumeRetryable: QuotaExceeded only.

export function isRetryable(e: unknown): boolean {
  return e instanceof UpstreamError || e instanceof QuotaExceeded
}

export function isResumeRetryable(e: unknown): boolean {
  return e instanceof QuotaExceeded
}

The SDK applies this for you. On the resume path it sets err.resumeSpent = !isResumeRetryable(err) and rethrows immediately when the token is spent, rather than retrying into a refusal that can only repeat. Read resumeSpent in your handler to decide between waiting and asking for approval again:

try {
  await fabric.resume('save_context', await loadToken(), { retry: true })
} catch (e) {
  if (e instanceof WexaError && e.resumeSpent) {
    // The approval is gone. Re-request it; do not loop.
  }
}

That branch is not theoretical. Sending a token the gateway would not redeem produced exactly this, live:

C resume on exhausted quota -> ApprovalError resumeSpent true

Where Python differs

The taxonomy, the retry membership and the resume boundary are the same. Four differences are worth knowing before you port a handler between the two.

TypeScriptPython
Nothing reached the gatewayTransportError, a WexaErrorurllib's OSError (a URLError), which is not a WexaError
The 504 classTimeoutErrorTimeoutError_, underscored to dodge the builtin
PredicatesisRetryable(e), isResumeRetryable(e)isinstance(e, RETRYABLE), isinstance(e, RESUME_RETRYABLE)
traceId on a 202CarriedDropped — a pending approval has no trace id

The Python README tells callers to catch WexaError and OSError; that advice is correct and necessary, because a one-line except WexaError there will not see a refused connection. In TypeScript one instanceof WexaError covers every failure, which is the reason TransportError was added.