> Source: https://wexa.ai/docs/administration/troubleshooting

# Troubleshooting

Most of what reaches an administrator arrives as "it doesn't work", and almost all of it is one of
a small number of things. This page is ordered by how often you will meet them, and every response
shape in it was executed rather than recalled.

## 1. "That tool isn't there"

This is the most common report and it is almost never a permissions problem. Every Wexa
instance carries all sixty-two tools, on the REST API, through both SDKs and on the Wexa MCP server
alike, whichever services sit behind it.

What varies is whether a tool can run. A tool whose backing service the deployment has not
configured is still listed, and a call to it fails with an error naming what is missing:
`<tool> is not configured (no data-service URL)`, `(no harness URL)` or
`(needs the harness and data-service URLs)`. It fails that way for everybody, including you, and
retrying will not help — the fix is the deployment's configuration.

### Diagnosing it in one request

`GET /v1/connection-info` is unauthenticated and returns the `tools` array the deployment
registered. A name missing from it is not a Wexa tool, usually a typo or a retired name. A name
present in it that answers with the not-configured error is a deployment that lacks the service
behind it.

## 2. "The tool is listed but it refuses when I call it"

A separate problem with a separate cause, and it looks like a bug because the tool was visibly
there a second ago.

**`tools/list` on the Wexa MCP server does not filter by the caller's grants.** Every registered
tool is listed to every authenticated caller. The grant check happens on the call, and its refusal
arrives inside a normal `200` response as an error result:

```
error (S2:resolve-scope): token missing required grant "…"
```

So "it was in the list" is not evidence the caller was allowed to use it. When someone reports this,
check what their credential actually carries with `GET /v1/whoami`, which returns the caller's
identifier, role, scope and grants. See
[credential management](/docs/administration/credential-management) for how grants are decided when
a key is minted, and [roles and permissions](/docs/administration/roles-and-permissions) for the
role behind them.

## 3. Telling the four refusals apart

When a developer sends you a failure, the status code is usually enough to route the problem — if
you know which four to distinguish. All of these were executed against one running gateway.

| What happened | Status | Body | Who fixes it |
|---|---|---|---|
| Permission denial | `403` | `{"error":"forbidden","error_description":"…"}` | You — it is a role or scope problem |
| Missing route | `404` | `404 page not found`, `text/plain` | The developer — the path is wrong |
| Retired endpoint | `410` | `{"error":"endpoint_moved", …}` | The developer — the path was real once |
| Wrong method | `405` | `Method Not Allowed`, `text/plain` | The developer — right path, wrong verb |
| No or bad credential | `401` | `{"error":"invalid_token"}` | The developer — or a revoked key |

### `403` — a permission denial

A JSON body with a sentence in it, and the sentence names the rule:

```json
403 {"error":"forbidden","error_description":"API key creation requires an admin role on this project"}
403 {"error":"forbidden","error_description":"changing project mode requires an admin role"}
403 {"error":"forbidden","error_description":"project outside token scope"}
```

**A `403` is the only one of the four you can fix.** The route exists, the method is right, the
credential is valid — the caller simply is not allowed. Read the description: "requires an admin
role" is a role problem, "outside token scope" is a project problem, and they have different fixes.

### `404` — a missing route

Go's own plain-text body, not a Wexa error envelope:

```
404 page not found
```

`text/plain`, no JSON, no `error` field. **If the body is not JSON, the gateway never recognised the
path at all** — nothing in Wexa produced that response, the router did. The path is misspelled,
or it belongs to a different service.

### `410` — a retired endpoint

The one that confuses people most, because it is the only failure that used to be a success.

```json
410 {"error":"endpoint_moved",
     "error_description":"This URL no longer works. Each project now has its own MCP connect URL …"}
```

Four paths return it: `POST`, `GET` and `DELETE` on `/mcp`, and `GET` on `/sse`. They are the
pre-per-project connect addresses. The replacement carries the project in the path —
`/mcp/{projectId}` — and the new address is on the Connect page for each project.

One trap inside the trap: **`POST /mcp/` with a trailing slash returns `404`, not `410`.** The
router will not match an empty project identifier, so it never reaches the retirement handler.
A developer who reports "sometimes 410, sometimes 404 on the same URL" has a trailing slash.

### `405` — a wrong method

```
405 Method Not Allowed
```

Plain text again. **The path is right and the verb is wrong** — which is useful, because it confirms
the route exists. `PUT /v1/apikeys` and `DELETE /v1/audit/verify` both produce it.

The shape of the API makes the right verb predictable: every tool is `POST /v1/<tool-name>` in
kebab-case, and everything else is conventional REST — `GET` for reads, plural collections,
sub-paths for actions.

### The one-question triage

Ask for the body, not just the code:

1. **Is the body JSON with an `error` field?** If not, the gateway's router answered, not Wexa:
   `404` means the path is unknown, `405` means the verb is wrong.
2. **Is `error` `endpoint_moved`?** The path is retired; move to `/mcp/{projectId}`.
3. **Is `error` `forbidden`?** Yours to fix — read the description for role versus scope.
4. **Is `error` `invalid_token`?** The credential, not the path.

## 4. "It worked this morning and now it's 429"

A `429` with a `Retry-After` header is a quota refusal, not a fault:

```
429 Retry-After: 24
{"error":"S3:rate-quota","error_description":"quota exceeded for project proj_d18 (retry after 24s)"}
```

Read the word before the identifier — `project` and `organization` are different problems. The
header is a real, computed number of seconds and a client that honours it recovers at the earliest
possible moment. Full treatment in
[usage and billing](/docs/administration/usage-and-billing).

## 5. "Saving context stopped working"

```json
422 {"error":"S4:validate-input",
     "error_description":"save-context is only available in Simple/auto mode; this project is in
       Advanced/manual mode — use create-ontology (proposal/commit)"}
```

Somebody changed the project's mode. This is the single most common consequence of that change, and
the refusal helpfully names the replacement. Ask for `GET /v1/projects/{projectId}/mode` and see
[project mode administration](/docs/administration/project-mode-administration).

## 6. "My change is stuck waiting for approval"

A consequential write made with an API key or an SDK session is held by default, and so is a
non-administrator's ontology commit in Advanced mode. The caller gets `"status":"pending_approval"`
with a resume token rather than a result. Nothing is wrong. Somebody has to decide it in the
Approvals Inbox, under **Actions → Approvals**. A held call is overdue after 48 hours and expires
unanswered after 96 hours.

If nobody should have to approve these writes, a project admin can turn the gate off in
**Settings → Projects**.

Three related reports have different causes:

- **`resume rejected: approval not approved (not approved, already used, or not yours)`** — the
  approval is still pending, was rejected or has expired, the resume token has already been spent,
  or it is being used by a different user, for a different tool or in a different project. Tokens
  are single-use; a retried resume always produces this.
- **`resume rejected: approval not found`** — the token is not a resume token at all.
- **`approval is already approved`**, `409` — two people decided the same item. The first decision
  stands.

See [approvals](/docs/administration/approval-workflows).

## 7. "Everything from yesterday is gone"

The gateway's audit trail at `GET /v1/audit`, the policy decision log and the OAuth clients are a live
view rather than an archive. They can be emptied, silently and without error. Approvals are stored,
and their events are kept in the platform audit trail under **Governance → Audit & Lineage**.

## What to collect before escalating

A report that contains these five things can be diagnosed without a second round trip:

1. **The full request** — method, path, and whether a credential was attached.
2. **The status code and the complete body**, not a paraphrase. The difference between JSON and
   plain text is half the diagnosis.
3. **The `lifecycle_id`** from the body if there is one. It joins the call to its audit events and
   its policy decision.
4. **`GET /v1/whoami`** from the same credential — role, scope and grants in one line.
5. **`GET /v1/connection-info`** from the same gateway — the tool count settles every
   "capability is missing" question immediately.
