Error handling and retries
Both SDKs ship a complete error hierarchy with retryability already decided. If you do not know it is there you will write a retry loop that is worse than the one the library already offers — one that retries a validation failure forever, or gives up on the one class of failure that was always going to succeed on the second attempt.
This page is written for the TypeScript SDK, but the taxonomy, the retry policy and the resume boundary are identical in Python. Where the two differ, the difference is named.
The hierarchy
Fourteen classes in TypeScript, thirteen in Python. Every one of them extends WexaError, so a
single catch can cover everything the library can throw.
WexaError base — carries stage, status, lifecycleId, traceId, raw, retryAfter
├── AuthError 401 invalid_token, unauthorized
├── ForbiddenError 403 the credential lacks the grant, or S2 refused the scope
│ └── PolicyDenied 403 S5 — a policy rule refused the call outright
├── ValidationError 400/422 malformed body, rejected Cypher, failed dry run (S4/S7)
├── ApprovalError 403 S6 — the resume token is bad, expired, rejected or spent
├── QuotaExceeded 429 S3 — over the request window
├── NotFound 404 no such approval, flow or execution in your scope
├── ConflictError 409 approval not pending, or agent not ready
├── UpstreamError 502 S8 — a downstream service refused or was unreachable
│ └── TimeoutError 504 an upstream deadline was exceeded
├── ConfigurationError 502/503 unconfigured, unavailable — a deployment fault
├── TransportError nothing reached the gateway (TypeScript only)
└── ApprovalRequired 202 a person must approve before the call proceeds
The inheritance is load-bearing rather than cosmetic. Python's retry policy is
isinstance(e, (UpstreamError, QuotaExceeded)), and the TypeScript port keeps exactly that
membership, so which class extends which decides what retries. TimeoutError retries because it
extends UpstreamError. ConfigurationError does not retry, despite carrying a 5xx status,
because it extends WexaError directly.
ApprovalRequired and ApprovalError are siblings, not parent and child. The first means a person
has not decided yet; the second means the token that would have resumed the call is no longer good.
What every error carries
| Field | What it is |
|---|---|
stage | The lifecycle stage that refused, S2:resolve-scope through S8:execute. undefined for refusals outside the lifecycle — a 401, a 404, unconfigured. |
status | HTTP status, or undefined when nothing reached the gateway. |
lifecycleId | The governed lifecycle this call ran as. A 401 predates it, so it is absent there. |
traceId | The value of the response's X-Trace-Id header. |
raw | The full decoded response body. On a TransportError it is the original thrown error instead. |
retryAfter | Seconds from a Retry-After header, clamped to [0, 60]. |
resumeSpent | Whether a resume token was consumed before this refusal. Only meaningful on the resume path. |
Quote lifecycleId and traceId in a support request. Both are appended to the error message as
well, so a bare log line already carries them:
quota exceeded for project proj_c10 (retry after 35s) [lifecycle_id=qlc_000157 trace_id=36088b71f728852753f230188485b0dd]
That line is the SDK's real output, not an illustration.
How a response becomes a class
raiseFor maps a non-2xx response in three tiers, first match wins.
- The
errorstring in the body. For a lifecycle refusal that string is the stage id, soS5:policybecomesPolicyDeniedandS3:rate-quotabecomesQuotaExceeded. The same table maps the non-lifecycle codes:invalid_tokenandunauthorizedtoAuthError,harness_unreachableanddata_service_unreachabletoUpstreamError,unconfiguredandunavailabletoConfigurationError. - The HTTP status. 400 and 422 to
ValidationError, 401 toAuthError, 403 toForbiddenError, 404 toNotFound, 409 toConflictError, 429 toQuotaExceeded, 504 toTimeoutError. - The fallback. Any 5xx becomes
UpstreamError; anything else becomes a bareWexaError.
502 and 503 are deliberately missing from tier 2. That omission is the whole reason an
unrecognised 502 is retryable while unconfigured is not: an unlabelled 502 falls through to the
fallback and becomes a retryable UpstreamError, whereas a 503 that names itself unconfigured is
caught in tier 1 and becomes a ConfigurationError that will never retry. Retrying a deployment
fault does not fix it; it just costs you three attempts before you tell an operator.
ApprovalRequired appears in none of these tables. A 202 is built in the transport layer, which is
the only place that holds the approval object and the tool name its constructor needs.
Catching them
Order the handlers from specific to general.
import {
ApprovalRequired, PolicyDenied, ForbiddenError, ValidationError,
QuotaExceeded, WexaError,
} from '@wexa-fabric/sdk'
try {
await fabric.saveContext({ nodes })
} catch (e) {
if (e instanceof ApprovalRequired) {
await checkpoint(e.approvalId, e.resumeToken) // durable, before you wait
} else if (e instanceof PolicyDenied) {
// S5 — a rule refused it. Terminal; a different credential will not help.
} else if (e instanceof ForbiddenError) {
// S2 — mint a credential that holds the grant.
} else if (e instanceof ValidationError) {
// Your payload. e.stage says whether the gateway or the SDK rejected it.
} else if (e instanceof QuotaExceeded) {
// S3 — wait e.retryAfter seconds, or let retry: true do it.
} else if (WexaError.isWexaError(e)) {
// Everything else the gateway sent, plus TransportError.
} else {
throw e
}
}
WexaError.isWexaError(e) is a string-tag check rather than a prototype check. Prefer it to
instanceof WexaError when two copies of the package could end up in one dependency tree:
instanceof fails across duplicated installs and this does not. It has no Python equivalent,
because Python cannot get two copies of a module into one interpreter the same way.
Which errors retry
Nothing retries unless you ask. The gateway has no idempotency key, so a retried runAgent starts a
second real run — the library refuses to make that decision for you.
await fabric.queryContext({ query }, { retry: true }) // safe: it is a read
retries: 3 on the constructor means three attempts in total, not three extra ones. Only
UpstreamError — and therefore TimeoutError — and QuotaExceeded are retried.
Every class in the hierarchy answers the two predicates like this:
| Class | isRetryable | isResumeRetryable |
|---|---|---|
WexaError | false | false |
AuthError | false | false |
ForbiddenError | false | false |
PolicyDenied | false | false |
ValidationError | false | false |
ApprovalError | false | false |
QuotaExceeded | true | true |
NotFound | false | false |
ConflictError | false | false |
UpstreamError | true | false |
TimeoutError | true | false |
ConfigurationError | false | false |
TransportError | false | false |
ApprovalRequired | false | false |
TransportError does not retry. Nothing reached the gateway, so the SDK cannot tell a wrong
workspace URL from a network blip, and it declines to guess.
Backoff
When the gateway sends Retry-After, that number is honoured — it is computed from the quota
window's real remaining time, and returning before the window resets only spends another 429. A
small random offset is added even then, so several clients told to wait the same number of seconds
do not all return in the same millisecond. With no header, backoff is full jitter — a uniform draw
from [0, 2^attempt) — capped at the 60-second quota window.
Retry-After is parsed strictly. RFC 9110 permits an HTTP-date as well as a count of seconds; the
gateway sends seconds, and the date form is ignored rather than misparsed into a wrong sleep. A
non-integer is ignored for the same reason.
Here is that path executed end to end against a live gateway, after deliberately exhausting the project's 120-request window:
A no retry -> QuotaExceeded status 429 stage S3:rate-quota retryAfter 35 isRetryable true isResumeRetryable true
B retry:true succeeded after 35.9 s; lifecycleId qlc_000159
Plain retry is not resume retry
This is the distinction that makes the two predicates worth having, and it is the one most likely to cost you a write.
A call that gates for approval throws ApprovalRequired carrying a single-use resumeToken.
When you send that token back, the gateway redeems it between S3 (quota) and S4 (validate).
S1 authenticate → S2 resolve-scope → S3 rate-quota → [REDEEM] → S4 validate-input → S5 policy → … → S8 execute
↑ ↑
refused here: token intact refused here: token gone
So a refusal at S3 arrives before the redemption and the token survives; every refusal after
that point means the token has been burned, the write did not land, and re-running needs a fresh
human approval. That is the entire content of isResumeRetryable: QuotaExceeded only.
export function isRetryable(e: unknown): boolean {
return e instanceof UpstreamError || e instanceof QuotaExceeded
}
export function isResumeRetryable(e: unknown): boolean {
return e instanceof QuotaExceeded
}
The SDK applies this for you. On the resume path it sets err.resumeSpent = !isResumeRetryable(err)
and rethrows immediately when the token is spent, rather than retrying into a refusal that can only
repeat. Read resumeSpent in your handler to decide between waiting and asking for approval again:
try {
await fabric.resume('save_context', await loadToken(), { retry: true })
} catch (e) {
if (e instanceof WexaError && e.resumeSpent) {
// The approval is gone. Re-request it; do not loop.
}
}
That branch is not theoretical. Sending a token the gateway would not redeem produced exactly this, live:
C resume on exhausted quota -> ApprovalError resumeSpent true
Where Python differs
The taxonomy, the retry membership and the resume boundary are the same. Four differences are worth knowing before you port a handler between the two.
| TypeScript | Python | |
|---|---|---|
| Nothing reached the gateway | TransportError, a WexaError | urllib's OSError (a URLError), which is not a WexaError |
| The 504 class | TimeoutError | TimeoutError_, underscored to dodge the builtin |
| Predicates | isRetryable(e), isResumeRetryable(e) | isinstance(e, RETRYABLE), isinstance(e, RESUME_RETRYABLE) |
traceId on a 202 | Carried | Dropped — a pending approval has no trace id |
The Python README tells callers to catch WexaError and OSError; that advice is correct and
necessary, because a one-line except WexaError there will not see a refused connection. In
TypeScript one instanceof WexaError covers every failure, which is the reason TransportError
was added.
Related
Calling tools, the non-tool methods, and cancellation.
The same failures as raw status codes and bodies, for callers with no SDK.
What makes a call gate in the first place.
Reading the lifecycle record a failed call left behind.