Troubleshooting
Most of what reaches an administrator arrives as "it doesn't work", and almost all of it is one of a small number of things. This page is ordered by how often you will meet them, and every response shape in it was executed rather than recalled.
1. "That tool isn't there"
This is the most common report and it is almost never a permissions problem. Every Wexa instance carries all sixty-two tools, on the REST API, through both SDKs and on the Wexa MCP server alike, whichever services sit behind it.
What varies is whether a tool can run. A tool whose backing service the deployment has not
configured is still listed, and a call to it fails with an error naming what is missing:
<tool> is not configured (no data-service URL), (no harness URL) or
(needs the harness and data-service URLs). It fails that way for everybody, including you, and
retrying will not help — the fix is the deployment's configuration.
Diagnosing it in one request
GET /v1/connection-info is unauthenticated and returns the tools array the deployment
registered. A name missing from it is not a Wexa tool, usually a typo or a retired name. A name
present in it that answers with the not-configured error is a deployment that lacks the service
behind it.
2. "The tool is listed but it refuses when I call it"
A separate problem with a separate cause, and it looks like a bug because the tool was visibly there a second ago.
tools/list on the Wexa MCP server does not filter by the caller's grants. Every registered
tool is listed to every authenticated caller. The grant check happens on the call, and its refusal
arrives inside a normal 200 response as an error result:
error (S2:resolve-scope): token missing required grant "…"
So "it was in the list" is not evidence the caller was allowed to use it. When someone reports this,
check what their credential actually carries with GET /v1/whoami, which returns the caller's
identifier, role, scope and grants. See
credential management for how grants are decided when
a key is minted, and roles and permissions for the
role behind them.
3. Telling the four refusals apart
When a developer sends you a failure, the status code is usually enough to route the problem — if you know which four to distinguish. All of these were executed against one running gateway.
| What happened | Status | Body | Who fixes it |
|---|---|---|---|
| Permission denial | 403 | {"error":"forbidden","error_description":"…"} | You — it is a role or scope problem |
| Missing route | 404 | 404 page not found, text/plain | The developer — the path is wrong |
| Retired endpoint | 410 | {"error":"endpoint_moved", …} | The developer — the path was real once |
| Wrong method | 405 | Method Not Allowed, text/plain | The developer — right path, wrong verb |
| No or bad credential | 401 | {"error":"invalid_token"} | The developer — or a revoked key |
403 — a permission denial
A JSON body with a sentence in it, and the sentence names the rule:
403 {"error":"forbidden","error_description":"API key creation requires an admin role on this project"}
403 {"error":"forbidden","error_description":"changing project mode requires an admin role"}
403 {"error":"forbidden","error_description":"project outside token scope"}
A 403 is the only one of the four you can fix. The route exists, the method is right, the
credential is valid — the caller simply is not allowed. Read the description: "requires an admin
role" is a role problem, "outside token scope" is a project problem, and they have different fixes.
404 — a missing route
Go's own plain-text body, not a Wexa error envelope:
404 page not found
text/plain, no JSON, no error field. If the body is not JSON, the gateway never recognised the
path at all — nothing in Wexa produced that response, the router did. The path is misspelled,
or it belongs to a different service.
410 — a retired endpoint
The one that confuses people most, because it is the only failure that used to be a success.
410 {"error":"endpoint_moved",
"error_description":"This URL no longer works. Each project now has its own MCP connect URL …"}
Four paths return it: POST, GET and DELETE on /mcp, and GET on /sse. They are the
pre-per-project connect addresses. The replacement carries the project in the path —
/mcp/{projectId} — and the new address is on the Connect page for each project.
One trap inside the trap: POST /mcp/ with a trailing slash returns 404, not 410. The
router will not match an empty project identifier, so it never reaches the retirement handler.
A developer who reports "sometimes 410, sometimes 404 on the same URL" has a trailing slash.
405 — a wrong method
405 Method Not Allowed
Plain text again. The path is right and the verb is wrong — which is useful, because it confirms
the route exists. PUT /v1/apikeys and DELETE /v1/audit/verify both produce it.
The shape of the API makes the right verb predictable: every tool is POST /v1/<tool-name> in
kebab-case, and everything else is conventional REST — GET for reads, plural collections,
sub-paths for actions.
The one-question triage
Ask for the body, not just the code:
- Is the body JSON with an
errorfield? If not, the gateway's router answered, not Wexa:404means the path is unknown,405means the verb is wrong. - Is
errorendpoint_moved? The path is retired; move to/mcp/{projectId}. - Is
errorforbidden? Yours to fix — read the description for role versus scope. - Is
errorinvalid_token? The credential, not the path.
4. "It worked this morning and now it's 429"
A 429 with a Retry-After header is a quota refusal, not a fault:
429 Retry-After: 24
{"error":"S3:rate-quota","error_description":"quota exceeded for project proj_d18 (retry after 24s)"}
Read the word before the identifier — project and organization are different problems. The
header is a real, computed number of seconds and a client that honours it recovers at the earliest
possible moment. Full treatment in
usage and billing.
5. "Saving context stopped working"
422 {"error":"S4:validate-input",
"error_description":"save-context is only available in Simple/auto mode; this project is in
Advanced/manual mode — use create-ontology (proposal/commit)"}
Somebody changed the project's mode. This is the single most common consequence of that change, and
the refusal helpfully names the replacement. Ask for GET /v1/projects/{projectId}/mode and see
project mode administration.
6. "My change is stuck waiting for approval"
A consequential write made with an API key or an SDK session is held by default, and so is a
non-administrator's ontology commit in Advanced mode. The caller gets "status":"pending_approval"
with a resume token rather than a result. Nothing is wrong. Somebody has to decide it in the
Approvals Inbox, under Actions → Approvals. A held call is overdue after 48 hours and expires
unanswered after 96 hours.
If nobody should have to approve these writes, a project admin can turn the gate off in Settings → Projects.
Three related reports have different causes:
resume rejected: approval not approved (not approved, already used, or not yours)— the approval is still pending, was rejected or has expired, the resume token has already been spent, or it is being used by a different user, for a different tool or in a different project. Tokens are single-use; a retried resume always produces this.resume rejected: approval not found— the token is not a resume token at all.approval is already approved,409— two people decided the same item. The first decision stands.
See approvals.
7. "Everything from yesterday is gone"
The gateway's audit trail at GET /v1/audit, the policy decision log and the OAuth clients are a live
view rather than an archive. They can be emptied, silently and without error. Approvals are stored,
and their events are kept in the platform audit trail under Governance → Audit & Lineage.
What to collect before escalating
A report that contains these five things can be diagnosed without a second round trip:
- The full request — method, path, and whether a credential was attached.
- The status code and the complete body, not a paraphrase. The difference between JSON and plain text is half the diagnosis.
- The
lifecycle_idfrom the body if there is one. It joins the call to its audit events and its policy decision. GET /v1/whoamifrom the same credential — role, scope and grants in one line.GET /v1/connection-infofrom the same gateway — the tool count settles every "capability is missing" question immediately.