On this page

Quota and credits

A quota is a ceiling on how much of a resource an organization may consume. Credits are the consumable balance that metered usage draws down.

These are different questions from whether an action is permitted, and Wexa answers them at a different point: quota is checked at stage S3:rate-quota of the lifecycle, before the tool has looked at the arguments and well before the policy engine is consulted. A perfectly permitted call can still be refused for want of quota, and a call that would have been denied by policy can be refused for quota first and never reach the policy engine at all.

What is enforced

Two ceilings are enforced on the governed call path, and one more on a single REST endpoint.

The project and organization request ceilings

Every governed tool call consumes one unit from two buckets at once: the project's and the organization's. Both are counted over a sliding window.

CeilingShipped defaultSetting
Requests per project per window120QUOTA_PER_PROJECT
Requests per organization per window600QUOTA_PER_ORG
Window length60 secondsdeployment configuration

A ceiling set to 0 means unlimited, and the bucket is not checked at all.

The window slides rather than resetting on the minute: the gateway prunes every timestamp older than the window before it counts, so a burst does not get a free reset at the top of the next minute.

The per-key ceiling on chat completions

The OpenAI-compatible /v1/chat/completions endpoint applies a separate ceiling of 60 requests per minute per API key, on top of the two above. This one uses fixed one-minute windows rather than a sliding one, and it applies only to that endpoint — it does not ceiling anything else a key can do.

How usage is counted

The counting rules are simple and worth stating exactly, because they decide what an operator can predict.

  • One governed tool call is one request, regardless of how much work the tool does. A query-context call returning one node and one returning fifty count the same.
  • Usage is counted per call, not per token, per node or per second of compute. Nothing in the quota path measures result size, graph writes or model usage.
  • Both buckets are consumed together, or neither is. If either bucket is already full the call is refused and no slot is consumed from either — so a caller hammering a full project bucket is not also burning the organization's.
  • A refused call does not consume a slot. Only a call that passes S3 is counted, which means a policy denial at S5 has already consumed one and an approval that parks at S6 has consumed one too.
  • A resumed call consumes another slot. Redemption of a resume token happens after S3, so the resume attempt is counted as its own request.
  • The surface makes no difference. A call over the Wexa MCP server, the REST API or either SDK is one request against the same two buckets.

The gateway records the usage that a call saw at the moment it passed — the project's count and ceiling, and the organization's — onto that call's trace, so a run that was slowed by a nearly-full bucket can be seen to have been.

What happens when a ceiling is exceeded

The call is refused at S3:rate-quota with 429, and the response carries a Retry-After header giving the number of seconds until the oldest timestamp in the offending bucket falls out of the window. It is a real number, not a fixed backoff: it is the earliest moment the call could succeed.

The refusal arrives as a tool result with isError set, so the model can read it and wait:

{
  "content": [
    {
      "type": "text",
      "text": "error (S3:rate-quota): quota exceeded for project prj_9f2a (retry after 37s)"
    }
  ],
  "isError": true
}

The retry hint is in the text. The Retry-After header is set on the REST response only.

Both SDKs treat quota exhaustion as retryable and wait for the number of seconds the header gave before trying again, rather than guessing. It is one of only two error classes either SDK will retry at all — the other being an upstream failure — and, as described on the policy and approvals page, it is the only one either will retry a resumed call through.

Where credits apply

Credits are a real part of the Wexa product: a balance is held against a workspace, top-ups are purchased, and consumption is reported in the product's own billing and utilization views.

They are not part of the governed call path. No stage of the lifecycle consults a credit balance, no tool call deducts from one, and no Wexa surface returns a credit balance or refuses a call for want of credits. The only thing that will refuse a call on grounds of volume is the quota described above.

Planning against the limits

Three practical consequences follow from the counting rules.

Batch where the tool lets you. Because one call is one request whatever its size, a single save-context carrying fifty nodes costs one unit and fifty calls carrying one node each cost fifty.

Expect an agent to spend more than its useful calls. A run that hits a policy denial or an approval gate has still consumed quota, and the resume consumes more. An agent doing exploratory work against a governed project should be sized with that in mind.

Watch the organization ceiling, not only the project one. The organization ceiling is shared. At the shipped defaults, five projects each running at their own full rate will exhaust it between them before any one of them reaches its own.

Where to look next