> Source: https://wexa.ai/docs/api/chat-completions

# Chat completions

`POST /v1/agents/{agentflowId}/chat/completions`

One route, and the only one on the gateway that speaks somebody else's protocol. Point an OpenAI
client at it and the call runs a governed Wexa agent underneath — the agent's own model, its own
tools, its own knowledge and its own guardrails — while the client sees the response shape it
expects.

## Pointing a client at it

The base URL is your gateway plus `/v1/agents/{agentflowId}`, and the credential goes in the usual
place.

```python
from openai import OpenAI

client = OpenAI(
    base_url="https://fabric.wexa.ai/v1/agents/af_abc123",
    api_key="fab_sk_…",
)

answer = client.chat.completions.create(
    model="ignored",
    messages=[{"role": "user", "content": "How did Q3 revenue go?"}],
)
print(answer.choices[0].message.content)
```

The raw call is just as short:

```bash
curl -sS -X POST https://fabric.wexa.ai/v1/agents/af_abc123/chat/completions \
  -H "Authorization: Bearer $FABRIC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o","messages":[{"role":"user","content":"How did Q3 revenue go?"}]}'
```

```json
{
  "id": "chatcmpl-exec_abc_77",
  "object": "chat.completion",
  "created": 1789469617,
  "model": "af_abc123",
  "choices": [{
    "index": 0,
    "message": { "role": "assistant", "content": "Q3 revenue was 4.2M, up 11% on Q2." },
    "finish_reason": "stop"
  }],
  "usage": { "prompt_tokens": 0, "completion_tokens": 812, "total_tokens": 812 }
}
```

`id` is `chatcmpl-` followed by the execution id, so a response can be traced back to the run that
produced it. `model` in the response is the agent id, not whatever you sent.

## Which request fields are real

Only two.

| Field | Effect |
|---|---|
| `messages` | the **last message with `role: "user"` and non-empty content** becomes the agent's goal |
| `stream` | switches the response to server-sent-event framing |

Everything else — `model`, `temperature`, `tools`, `top_p`, `max_tokens`, `response_format` — is
accepted and ignored. The call above sent `"model":"gpt-4o"`, `"temperature":0.2` and a `tools`
array, and the execution target received only
`{"goal":"How did Q3 revenue go?","caller_id":"user_abc","project_id":"proj_abc"}`.

Earlier messages are not discarded from the record, but only the last user message becomes the goal.
A system prompt sent alongside it does not reshape the agent. If there is no user message at all:

```json
{ "error": { "message": "messages[] must include at least one message with role \"user\"",
             "type": "invalid_request_error", "code": "missing_messages" } }
```

## Streaming is framing, not incremental tokens

`"stream": true` produces exactly the frames an OpenAI client expects, in the right order:

```text
data: {"id":"chatcmpl-exec_abc_77","object":"chat.completion.chunk","created":1789469617,"model":"af_abc123","choices":[{"index":0,"delta":{"role":"assistant","content":"Q3 revenue was 4.2M, up 11% on Q2."},"finish_reason":null}]}

data: {"id":"chatcmpl-exec_abc_77","object":"chat.completion.chunk","created":1789469617,"model":"af_abc123","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]
```

`Content-Type` becomes `text/event-stream`, and the whole answer arrives in the first chunk.

## Rate limit: 60 per minute, per key

This route is the **only** place on the gateway where the per-key limiter is consulted.

A single fresh key was fired at the endpoint 65 times in one window. The first 60 answered `200`;
the 61st and everything after it answered:

```text
HTTP/1.1 429 Too Many Requests
Retry-After: 1

{"error":{"message":"too many requests for this API key — try again shortly","type":"rate_limit_exceeded","code":"rate_limited"}}
```

A second key, minted on the same project, answered `200` at the same moment — the window is per key
id, not per project or per user.

Three details that change how you should retry:

- The window is a **fixed** one minute, not a sliding one. The allowance resets wholesale rather
  than draining.
- `Retry-After` is a hard-coded `1`, not the real seconds remaining. Back off on your own schedule;
  waiting one second immediately after the 61st call will usually be refused again.

The project and organization quota windows are separate, apply to the governed tool routes, and
produce Wexa's own error body with a genuinely computed `Retry-After`.
[Quota and credits](/docs/concepts/quota-and-credits) covers them.

A credential with no key id — a user token — is not subject to this limiter at all. That was
confirmed: the same request with a user token answered `200`.

## Every failure shape

Seven of the ten were produced live; the ledger at the foot of the page says which.

| Status | `type` / `code` | When |
|---|---|---|
| `400` | `invalid_request_error` / `invalid_json` | the body is not JSON |
| `400` | `invalid_request_error` / `missing_messages` | no message with `role: "user"` |
| `403` | `insufficient_scope` / `forbidden` | the credential lacks the `agent:run` grant |
| `403` | `insufficient_scope` / `forbidden` | the agent belongs to another project |
| `404` | `invalid_request_error` / `agent_not_found` | no such agent |
| `409` | `invalid_request_error` / `not_ready` | the agent has no promoted version |
| `429` | `rate_limit_exceeded` / `rate_limited` | the per-key limit above |
| `502` | `upstream_error` / `data_service_unreachable` | the execution target is unreachable |
| `503` | `unavailable` / `not_configured` | the deployment has no execution target configured |
| `504` | `timeout` / `execution_timeout` | the run exceeded the deployment's timeout |

The `409` deserves a note, because its message is unusually long on purpose:

```json
{ "error": { "type": "invalid_request_error", "code": "not_ready",
  "message": "agentflow has no promoted version promote-flow will NOT fix this: it snapshots the flow's shape, not the agent's config. If this agent already has versions, rollback-agent-version makes one live and this endpoint then works — list them with agent-versions. Creating the FIRST version needs the Fabric console (Agent Simulation, an Enterprise-tier feature). run-process-flow and run-agent require no promotion at all." } }
```

An agent must have a promoted version before this endpoint will run it. If yours does not,
`run-agent` and `run-process-flow` need no promotion and are the way through.

## Four compatibility limits

1. **`401` does not speak OpenAI.** Every other failure on this route uses the OpenAI error
   envelope; a missing or invalid credential is refused by the gateway's shared authentication layer
   *before* the handler runs, and comes back as Wexa's own shape —
   `{"error":"invalid_token","error_description":"missing bearer credential"}`. An OpenAI client that
   parses errors strictly will fail to parse that one. Handle `401` separately.
2. **`usage.prompt_tokens` is always `0`.** The execution target reports a total; the gateway puts
   it in both `total_tokens` and `completion_tokens` and leaves the prompt count at zero. Do not
   compute cost from these fields.
3. **`n`, `logprobs`, `tools` and function calling are not supported.** There is always exactly one
   choice, `finish_reason` is always `stop` on success, and tool use happens inside the agent rather
   than being negotiated with the caller.
4. **There is no `GET /v1/models`.** The agent id in the URL is the model selector; a client that
   lists models at startup will not find one.

## Per-route ledger

| Case | Result |
|---|---|
| happy path, extra OpenAI fields sent | `200`; execution target received only `goal`, `caller_id`, `project_id` |
| `stream: true` | `200`, `text/event-stream`, three frames as quoted |
| no user message | `400 missing_messages` |
| no promoted version | `409 not_ready` with the full remedy message |
| unknown agent | `404 agent_not_found` |
| agent in another project | `403 forbidden` |
| key without `agent:run` | `403 forbidden` |
| no credential | `401`, in Wexa's envelope, not OpenAI's |
| user token instead of a key | `200`, not rate-limited |
| 65 calls on one key | 60 × `200`, then `429` with `Retry-After: 1` |
| second key during the refusal | `200` |
| execution target unreachable | `502 data_service_unreachable` |

The status mapping, the request translation and the rate limit are gateway behaviour. The answer text and the token count come from whatever execution target your deployment runs.
