Usage and billing
An administrator asked "what is this costing us, and what happens when we hit a limit" needs two separate answers, because Wexa has two separate mechanisms and they do not meet. This page gives both, and is explicit about the join between them that does not exist.
The two mechanisms
Quota is a request ceiling. It is enforced by the gateway, in front of every governed call, and it refuses. It has nothing to do with money.
Credits are a balance. They are real, and they are held and spent in the main application, not service. They are not enforced on the governed call path.
How usage is measured
Every governed call consumes exactly one unit from two buckets at once, before it executes.
| Bucket | Default ceiling | Window |
|---|---|---|
| Per project | 120 requests | sliding 60 seconds |
| Per organization | 600 requests | sliding 60 seconds |
The window slides: the counter is the number of calls in the last sixty seconds, not calls since a
clock minute began. A ceiling set to 0 means unlimited for that bucket.
Three properties are worth knowing precisely, because each one changes how you reason about a noisy project:
- Both buckets are consumed, or neither is. A call refused by the project ceiling does not spend an organization unit, so one runaway project cannot drain the organization's budget through its own refusals.
- A refused call consumes nothing at all. A client retrying into a
429in a tight loop does not push its own recovery further away. The retry is free; only success costs. - A resumed call consumes another unit. A call held for approval spent a unit when it was first made and spends a second one when it is resumed. Budget for gated operations at twice the request count.
You can read your own consumption at any time. The Wexa documentation tool's governance topic returns it:
curl -X POST $FABRIC_URL/v1/docs -H "Authorization: Bearer $TOKEN" \
-d '{"topic":"governance"}'
"quota": { "project_window_used": 114, "project_window_max": 120,
"org_window_used": 114, "org_window_max": 600 }
That is the cheapest way to answer "how close are we", and it is available to any caller.
What happens when a ceiling is reached
The call is refused. This is a real refusal, captured as it came back on the 121st call inside the window:
HTTP/1.1 429 Too Many Requests
Retry-After: 24
{"error":"S3:rate-quota",
"error_description":"quota exceeded for project proj_d18 (retry after 24s)",
"lifecycle_id":"qlc_000126"}
Three things to point out to whoever brings you this error.
The message names the bucket. for project … and for organization … are different problems:
one is a single noisy project, the other is every project at once. Read the word before the
identifier.
Retry-After is computed, not a constant. It is the time until the oldest call in the window
ages out. Sampled over one recovery, it read 24, then 6, and then the call succeeded — so a
client that honours the header recovers at the earliest possible moment, and one that ignores it
just burns refusals. Both SDKs honour it; see
errors and retries.
The refusal is recorded. It lands in the audit trail as a rejection at stage S3:rate-quota,
so a report of "we were throttled at 14:05" is checkable — for as long as the trail survives. See
audit and compliance.
The separate per-credential ceiling
One more limit exists and is easy to mistake for the others: the OpenAI-compatible chat-completions endpoint allows 60 requests a minute per credential, in fixed one-minute windows. It applies to that one endpoint and nothing else, it is per key rather than per project, and it sits on top of the project and organization ceilings rather than replacing them.
Where credits actually live
Credits are a real product surface — just not this one. They are held and spent through the
/credits/* endpoints, reached from the main application rather than through any of the four
surfaces. Nothing in the governed call path reads or writes them.
What this means for you as an administrator:
- The credit balance is the answer to "what did we buy". The quota counters are the answer to "how hard are we calling". Neither answers the other.
- A drained balance does not stop governed calls, and exhausted quota does not preserve credit. Do not use one as a safety net for the other.
- The screens that show and top up a balance are in the main application, under billing, which is an owner-tier capability. See roles and permissions for who reaches it.
Reconciling a bill
Given the above, reconciliation is possible but partial, and it is better to know the shape of the gap before the invoice arrives than afterwards.
What Wexa can corroborate:
- That calls happened, and which. Every governed call writes an audit event naming the tool,
the actor, the project and the lifecycle identifier. Counting
*.executedevents per project over a period is a defensible usage figure. - That refusals happened. Rejections at
S3:rate-quotaare recorded separately from executions, so throttled traffic is not counted as served traffic. - Which project each call belonged to — for reading, though not as a cryptographic assurance; the project identifier is recorded but is outside the audit hash chain.
What it cannot:
- Anything predating the last restart. The audit trail is not durable. Unless you are already archiving it, the usage record for last month does not exist. This is the single largest obstacle to reconciling a bill from Wexa's own data, and the archiving routine in audit and compliance is the fix.
- A per-call cost. No cost, token count or credit amount is attached to any event on the governed path.
- A tie between a call and a credit movement. There is no identifier shared between the audit trail and the credits ledger, so the two cannot be joined even by hand.
A workable routine
- Archive
GET /v1/auditcontinuously into your own storage. Without this, nothing below is possible. - Count
*.executedevents per project per period from the archive. That is your usage denominator. - Take the credit ledger from the main application as your cost numerator.
- Reconcile at the project level, not the call level, and say so in the reconciliation. The per-call join does not exist, and a report that implies it does will not survive a question.
Related reading
Quota and credits is the concept treatment of the same two mechanisms. Credential management covers the per-key ceiling and who should hold a key at all.