On this page

Troubleshooting

Most of what reaches an administrator arrives as "it doesn't work", and almost all of it is one of a small number of things. This page is ordered by how often you will meet them, and every response shape in it was executed rather than recalled.

1. "That tool isn't there"

This is the most common report and it is almost never a permissions problem. Every Wexa instance carries all sixty-two tools, on the REST API, through both SDKs and on the Wexa MCP server alike, whichever services sit behind it.

What varies is whether a tool can run. A tool whose backing service the deployment has not configured is still listed, and a call to it fails with an error naming what is missing: <tool> is not configured (no data-service URL), (no harness URL) or (needs the harness and data-service URLs). It fails that way for everybody, including you, and retrying will not help — the fix is the deployment's configuration.

Diagnosing it in one request

GET /v1/connection-info is unauthenticated and returns the tools array the deployment registered. A name missing from it is not a Wexa tool, usually a typo or a retired name. A name present in it that answers with the not-configured error is a deployment that lacks the service behind it.

2. "The tool is listed but it refuses when I call it"

A separate problem with a separate cause, and it looks like a bug because the tool was visibly there a second ago.

tools/list on the Wexa MCP server does not filter by the caller's grants. Every registered tool is listed to every authenticated caller. The grant check happens on the call, and its refusal arrives inside a normal 200 response as an error result:

error (S2:resolve-scope): token missing required grant "…"

So "it was in the list" is not evidence the caller was allowed to use it. When someone reports this, check what their credential actually carries with GET /v1/whoami, which returns the caller's identifier, role, scope and grants. See credential management for how grants are decided when a key is minted, and roles and permissions for the role behind them.

3. Telling the four refusals apart

When a developer sends you a failure, the status code is usually enough to route the problem — if you know which four to distinguish. All of these were executed against one running gateway.

What happenedStatusBodyWho fixes it
Permission denial403{"error":"forbidden","error_description":"…"}You — it is a role or scope problem
Missing route404404 page not found, text/plainThe developer — the path is wrong
Retired endpoint410{"error":"endpoint_moved", …}The developer — the path was real once
Wrong method405Method Not Allowed, text/plainThe developer — right path, wrong verb
No or bad credential401{"error":"invalid_token"}The developer — or a revoked key

403 — a permission denial

A JSON body with a sentence in it, and the sentence names the rule:

403 {"error":"forbidden","error_description":"API key creation requires an admin role on this project"}
403 {"error":"forbidden","error_description":"changing project mode requires an admin role"}
403 {"error":"forbidden","error_description":"project outside token scope"}

A 403 is the only one of the four you can fix. The route exists, the method is right, the credential is valid — the caller simply is not allowed. Read the description: "requires an admin role" is a role problem, "outside token scope" is a project problem, and they have different fixes.

404 — a missing route

Go's own plain-text body, not a Wexa error envelope:

404 page not found

text/plain, no JSON, no error field. If the body is not JSON, the gateway never recognised the path at all — nothing in Wexa produced that response, the router did. The path is misspelled, or it belongs to a different service.

410 — a retired endpoint

The one that confuses people most, because it is the only failure that used to be a success.

410 {"error":"endpoint_moved",
     "error_description":"This URL no longer works. Each project now has its own MCP connect URL …"}

Four paths return it: POST, GET and DELETE on /mcp, and GET on /sse. They are the pre-per-project connect addresses. The replacement carries the project in the path — /mcp/{projectId} — and the new address is on the Connect page for each project.

One trap inside the trap: POST /mcp/ with a trailing slash returns 404, not 410. The router will not match an empty project identifier, so it never reaches the retirement handler. A developer who reports "sometimes 410, sometimes 404 on the same URL" has a trailing slash.

405 — a wrong method

405 Method Not Allowed

Plain text again. The path is right and the verb is wrong — which is useful, because it confirms the route exists. PUT /v1/apikeys and DELETE /v1/audit/verify both produce it.

The shape of the API makes the right verb predictable: every tool is POST /v1/<tool-name> in kebab-case, and everything else is conventional REST — GET for reads, plural collections, sub-paths for actions.

The one-question triage

Ask for the body, not just the code:

  1. Is the body JSON with an error field? If not, the gateway's router answered, not Wexa: 404 means the path is unknown, 405 means the verb is wrong.
  2. Is error endpoint_moved? The path is retired; move to /mcp/{projectId}.
  3. Is error forbidden? Yours to fix — read the description for role versus scope.
  4. Is error invalid_token? The credential, not the path.

4. "It worked this morning and now it's 429"

A 429 with a Retry-After header is a quota refusal, not a fault:

429 Retry-After: 24
{"error":"S3:rate-quota","error_description":"quota exceeded for project proj_d18 (retry after 24s)"}

Read the word before the identifier — project and organization are different problems. The header is a real, computed number of seconds and a client that honours it recovers at the earliest possible moment. Full treatment in usage and billing.

5. "Saving context stopped working"

422 {"error":"S4:validate-input",
     "error_description":"save-context is only available in Simple/auto mode; this project is in
       Advanced/manual mode — use create-ontology (proposal/commit)"}

Somebody changed the project's mode. This is the single most common consequence of that change, and the refusal helpfully names the replacement. Ask for GET /v1/projects/{projectId}/mode and see project mode administration.

6. "My change is stuck waiting for approval"

A consequential write made with an API key or an SDK session is held by default, and so is a non-administrator's ontology commit in Advanced mode. The caller gets "status":"pending_approval" with a resume token rather than a result. Nothing is wrong. Somebody has to decide it in the Approvals Inbox, under Actions → Approvals. A held call is overdue after 48 hours and expires unanswered after 96 hours.

If nobody should have to approve these writes, a project admin can turn the gate off in Settings → Projects.

Three related reports have different causes:

  • resume rejected: approval not approved (not approved, already used, or not yours) — the approval is still pending, was rejected or has expired, the resume token has already been spent, or it is being used by a different user, for a different tool or in a different project. Tokens are single-use; a retried resume always produces this.
  • resume rejected: approval not found — the token is not a resume token at all.
  • approval is already approved, 409 — two people decided the same item. The first decision stands.

See approvals.

7. "Everything from yesterday is gone"

The gateway's audit trail at GET /v1/audit, the policy decision log and the OAuth clients are a live view rather than an archive. They can be emptied, silently and without error. Approvals are stored, and their events are kept in the platform audit trail under Governance → Audit & Lineage.

What to collect before escalating

A report that contains these five things can be diagnosed without a second round trip:

  1. The full request — method, path, and whether a credential was attached.
  2. The status code and the complete body, not a paraphrase. The difference between JSON and plain text is half the diagnosis.
  3. The lifecycle_id from the body if there is one. It joins the call to its audit events and its policy decision.
  4. GET /v1/whoami from the same credential — role, scope and grants in one line.
  5. GET /v1/connection-info from the same gateway — the tool count settles every "capability is missing" question immediately.