> Source: https://wexa.ai/docs/surfaces/typescript/error-handling

# Error handling and retries

Both SDKs ship a complete error hierarchy with retryability already decided. If you do not know it
is there you will write a retry loop that is worse than the one the library already offers — one
that retries a validation failure forever, or gives up on the one class of failure that was always
going to succeed on the second attempt.

This page is written for the TypeScript SDK, but the taxonomy, the retry policy and the resume
boundary are identical in Python. Where the two differ, the difference is named.

## The hierarchy

Fourteen classes in TypeScript, thirteen in Python. Every one of them extends `WexaError`, so a
single `catch` can cover everything the library can throw.

```
WexaError                     base — carries stage, status, lifecycleId, traceId, raw, retryAfter
├── AuthError                 401  invalid_token, unauthorized
├── ForbiddenError            403  the credential lacks the grant, or S2 refused the scope
│   └── PolicyDenied          403  S5 — a policy rule refused the call outright
├── ValidationError           400/422  malformed body, rejected Cypher, failed dry run (S4/S7)
├── ApprovalError             403  S6 — the resume token is bad, expired, rejected or spent
├── QuotaExceeded             429  S3 — over the request window
├── NotFound                  404  no such approval, flow or execution in your scope
├── ConflictError             409  approval not pending, or agent not ready
├── UpstreamError             502  S8 — a downstream service refused or was unreachable
│   └── TimeoutError          504  an upstream deadline was exceeded
├── ConfigurationError        502/503  unconfigured, unavailable — a deployment fault
├── TransportError            nothing reached the gateway (TypeScript only)
└── ApprovalRequired          202  a person must approve before the call proceeds
```

The inheritance is load-bearing rather than cosmetic. Python's retry policy is
`isinstance(e, (UpstreamError, QuotaExceeded))`, and the TypeScript port keeps exactly that
membership, so **which class extends which decides what retries**. `TimeoutError` retries because it
extends `UpstreamError`. `ConfigurationError` does *not* retry, despite carrying a 5xx status,
because it extends `WexaError` directly.

`ApprovalRequired` and `ApprovalError` are siblings, not parent and child. The first means a person
has not decided yet; the second means the token that would have resumed the call is no longer good.

## What every error carries

| Field | What it is |
|---|---|
| `stage` | The lifecycle stage that refused, `S2:resolve-scope` through `S8:execute`. `undefined` for refusals outside the lifecycle — a 401, a 404, `unconfigured`. |
| `status` | HTTP status, or `undefined` when nothing reached the gateway. |
| `lifecycleId` | The governed lifecycle this call ran as. A 401 predates it, so it is absent there. |
| `traceId` | The value of the response's `X-Trace-Id` header. |
| `raw` | The full decoded response body. On a `TransportError` it is the original thrown error instead. |
| `retryAfter` | Seconds from a `Retry-After` header, clamped to `[0, 60]`. |
| `resumeSpent` | Whether a resume token was consumed before this refusal. Only meaningful on the resume path. |

Quote `lifecycleId` and `traceId` in a support request. Both are appended to the error message as
well, so a bare log line already carries them:

```text
quota exceeded for project proj_c10 (retry after 35s) [lifecycle_id=qlc_000157 trace_id=36088b71f728852753f230188485b0dd]
```

That line is the SDK's real output, not an illustration.

## How a response becomes a class

`raiseFor` maps a non-2xx response in three tiers, first match wins.

1. **The `error` string in the body.** For a lifecycle refusal that string *is* the stage id, so
   `S5:policy` becomes `PolicyDenied` and `S3:rate-quota` becomes `QuotaExceeded`. The same table
   maps the non-lifecycle codes: `invalid_token` and `unauthorized` to `AuthError`,
   `harness_unreachable` and `data_service_unreachable` to `UpstreamError`, `unconfigured` and
   `unavailable` to `ConfigurationError`.
2. **The HTTP status.** 400 and 422 to `ValidationError`, 401 to `AuthError`, 403 to
   `ForbiddenError`, 404 to `NotFound`, 409 to `ConflictError`, 429 to `QuotaExceeded`, 504 to
   `TimeoutError`.
3. **The fallback.** Any 5xx becomes `UpstreamError`; anything else becomes a bare `WexaError`.

**502 and 503 are deliberately missing from tier 2.** That omission is the whole reason an
unrecognised 502 is retryable while `unconfigured` is not: an unlabelled 502 falls through to the
fallback and becomes a retryable `UpstreamError`, whereas a 503 that names itself `unconfigured` is
caught in tier 1 and becomes a `ConfigurationError` that will never retry. Retrying a deployment
fault does not fix it; it just costs you three attempts before you tell an operator.

`ApprovalRequired` appears in none of these tables. A 202 is built in the transport layer, which is
the only place that holds the approval object and the tool name its constructor needs.

## Catching them

Order the handlers from specific to general.

```ts
import {
  ApprovalRequired, PolicyDenied, ForbiddenError, ValidationError,
  QuotaExceeded, WexaError,
} from '@wexa-fabric/sdk'

try {
  await fabric.saveContext({ nodes })
} catch (e) {
  if (e instanceof ApprovalRequired) {
    await checkpoint(e.approvalId, e.resumeToken)   // durable, before you wait
  } else if (e instanceof PolicyDenied) {
    // S5 — a rule refused it. Terminal; a different credential will not help.
  } else if (e instanceof ForbiddenError) {
    // S2 — mint a credential that holds the grant.
  } else if (e instanceof ValidationError) {
    // Your payload. e.stage says whether the gateway or the SDK rejected it.
  } else if (e instanceof QuotaExceeded) {
    // S3 — wait e.retryAfter seconds, or let retry: true do it.
  } else if (WexaError.isWexaError(e)) {
    // Everything else the gateway sent, plus TransportError.
  } else {
    throw e
  }
}
```

`WexaError.isWexaError(e)` is a string-tag check rather than a prototype check. Prefer it to
`instanceof WexaError` when two copies of the package could end up in one dependency tree:
`instanceof` fails across duplicated installs and this does not. It has no Python equivalent,
because Python cannot get two copies of a module into one interpreter the same way.

## Which errors retry

Nothing retries unless you ask. The gateway has no idempotency key, so a retried `runAgent` starts a
second real run — the library refuses to make that decision for you.

```ts
await fabric.queryContext({ query }, { retry: true })   // safe: it is a read
```

`retries: 3` on the constructor means **three attempts in total**, not three extra ones. Only
`UpstreamError` — and therefore `TimeoutError` — and `QuotaExceeded` are retried.

Every class in the hierarchy answers the two predicates like this:

| Class | `isRetryable` | `isResumeRetryable` |
|---|---|---|
| `WexaError` | false | false |
| `AuthError` | false | false |
| `ForbiddenError` | false | false |
| `PolicyDenied` | false | false |
| `ValidationError` | false | false |
| `ApprovalError` | false | false |
| `QuotaExceeded` | **true** | **true** |
| `NotFound` | false | false |
| `ConflictError` | false | false |
| `UpstreamError` | **true** | false |
| `TimeoutError` | **true** | false |
| `ConfigurationError` | false | false |
| `TransportError` | false | false |
| `ApprovalRequired` | false | false |

`TransportError` does not retry. Nothing reached the gateway, so the SDK cannot tell a wrong
workspace URL from a network blip, and it declines to guess.

### Backoff

When the gateway sends `Retry-After`, that number is honoured — it is computed from the quota
window's real remaining time, and returning before the window resets only spends another 429. A
small random offset is added even then, so several clients told to wait the same number of seconds
do not all return in the same millisecond. With no header, backoff is full jitter — a uniform draw
from `[0, 2^attempt)` — capped at the 60-second quota window.

`Retry-After` is parsed strictly. RFC 9110 permits an HTTP-date as well as a count of seconds; the
gateway sends seconds, and the date form is ignored rather than misparsed into a wrong sleep. A
non-integer is ignored for the same reason.

Here is that path executed end to end against a live gateway, after deliberately exhausting the
project's 120-request window:

```text
A no retry -> QuotaExceeded status 429 stage S3:rate-quota retryAfter 35 isRetryable true isResumeRetryable true
B retry:true succeeded after 35.9 s; lifecycleId qlc_000159
```

## Plain retry is not resume retry

This is the distinction that makes the two predicates worth having, and it is the one most likely to
cost you a write.

A call that gates for approval throws `ApprovalRequired` carrying a **single-use** `resumeToken`.
When you send that token back, the gateway redeems it **between S3 (quota) and S4 (validate)**.

```
S1 authenticate → S2 resolve-scope → S3 rate-quota → [REDEEM] → S4 validate-input → S5 policy → … → S8 execute
                                          ↑                          ↑
                          refused here: token intact      refused here: token gone
```

So a refusal at S3 arrives before the redemption and the token survives; **every** refusal after
that point means the token has been burned, the write did not land, and re-running needs a fresh
human approval. That is the entire content of `isResumeRetryable`: `QuotaExceeded` only.

```ts
export function isRetryable(e: unknown): boolean {
  return e instanceof UpstreamError || e instanceof QuotaExceeded
}

export function isResumeRetryable(e: unknown): boolean {
  return e instanceof QuotaExceeded
}
```

The SDK applies this for you. On the resume path it sets `err.resumeSpent = !isResumeRetryable(err)`
and rethrows immediately when the token is spent, rather than retrying into a refusal that can only
repeat. Read `resumeSpent` in your handler to decide between waiting and asking for approval again:

```ts
try {
  await fabric.resume('save_context', await loadToken(), { retry: true })
} catch (e) {
  if (e instanceof WexaError && e.resumeSpent) {
    // The approval is gone. Re-request it; do not loop.
  }
}
```

That branch is not theoretical. Sending a token the gateway would not redeem produced exactly this,
live:

```text
C resume on exhausted quota -> ApprovalError resumeSpent true
```

## Where Python differs

The taxonomy, the retry membership and the resume boundary are the same. Four differences are worth
knowing before you port a handler between the two.

| | TypeScript | Python |
|---|---|---|
| Nothing reached the gateway | `TransportError`, a `WexaError` | urllib's `OSError` (a `URLError`), which is **not** a `WexaError` |
| The 504 class | `TimeoutError` | `TimeoutError_`, underscored to dodge the builtin |
| Predicates | `isRetryable(e)`, `isResumeRetryable(e)` | `isinstance(e, RETRYABLE)`, `isinstance(e, RESUME_RETRYABLE)` |
| `traceId` on a 202 | Carried | Dropped — a pending approval has no trace id |

The Python README tells callers to catch `WexaError` and `OSError`; that advice is correct and
necessary, because a one-line `except WexaError` there will not see a refused connection. In
TypeScript one `instanceof WexaError` covers every failure, which is the reason `TransportError`
was added.

## Related

  <Card title="TypeScript SDK — usage" href="/docs/surfaces/typescript/usage">
    Calling tools, the non-tool methods, and cancellation.
  </Card>
  <Card title="REST errors and retries" href="/docs/surfaces/rest/errors-and-retries">
    The same failures as raw status codes and bodies, for callers with no SDK.
  </Card>
  <Card title="Policy and approvals" href="/docs/concepts/policy-and-approvals">
    What makes a call gate in the first place.
  </Card>
  <Card title="Lifecycle and audit" href="/docs/concepts/lifecycle-and-audit">
    Reading the lifecycle record a failed call left behind.
  </Card>
