> Source: https://wexa.ai/docs/concepts/models

# Models

A **model** is the LLM an agent uses: either one the platform provides or one you bring. Every
[agent](/docs/concepts/agents) needs one, and the choice has two consequences — what the agent is
capable of, and who pays for it.

## How an agent gets its model

An agent's definition carries an `llm` block whose `model` names the LLM. It defaults to the
sentinel value `system_model`, which does not name a model at all: it means *whichever model this
organization has chosen*.

That gives two levers, at two different levels.

- **Per agent.** Put a real model id in that agent's `llm.model`. Only that agent is affected, and
  it stops following the organization default.
- **Per organization.** [`set-model`](/docs/tools/set-model) changes the default itself. Every agent
  still using `system_model` follows it immediately, across every project in the organization.
  Agents that name their own model are unaffected.

Because the second reaches every project, it requires an admin role. It also takes an optional
`fallback` — a second model used when the default is unavailable.

## Finding out which models you can use

[`list-models`](/docs/tools/list-models) takes no arguments — which organization is being asked
about comes from your credential — and returns the models this organization can run, along with
which one is the default.

It answers with two families of id, and both are valid in an agent's `llm.model`.

- The **registry** entries are the organization's own model registry: an id, a name, a provider, a
  health status, and whether it is the default. These are the ones to prefer.
- The **available** list is what the running service reports it can reach: the `system_model`
  sentinel, plus composite ids identifying a particular model from a particular provider for a
  particular organization.

The registry is read from a separate service, so it is possible for `list-models` to answer while
that part is unavailable. When that happens the result says so rather than returning a short list
that looks complete. Treat a missing registry as "not known right now", not as "this organization
has no registry models".

`set-model` refuses a model id the organization does not actually have configured, and refuses an
unconfigured `fallback` too. That refusal is doing you a favour: an accepted id that resolves to
nothing would break every agent in the organization on its next run, and a rejected call is a much
cheaper way to find out.

## Calling a model directly

Everything above is about which model an *agent* uses. You can also call a model yourself, with no
agent involved: [`run-model`](/docs/tools/run-model) takes a model and a prompt and returns the
answer, the tokens it consumed, what it cost and how long it took.

### What you pass

`model` accepts any id `list-models` returns — a registry id or a composite — and is **optional**.
Omit it and the call runs the organization's default, resolved exactly as it is for an agent on
`system_model`. A free-text provider name such as `gpt-4o` is refused: a model has to be one this
organization has registered, so that every call resolves to a real credential and a permitted host.

Send either a `prompt` — the simple form, treated as one user message — or a `messages` list of
`role`/`content` pairs for a system message and a multi-turn conversation. When both are present,
`messages` wins. The roles are `system`, `user` and `assistant`; there is no tool role, because tool
calling is not offered here.

These are the controls you may set, and they are an allow-list rather than a pass-through:

| Field | Meaning | Default |
| --- | --- | --- |
| `max_tokens` | Ceiling on the tokens generated | 512 |
| `temperature` | 0 to 2 | 0.2 |
| `top_p` | Nucleus sampling, 0 to 1 | provider's |
| `stop` | Up to four sequences that end generation | none |
| `seed` | Best-effort determinism, where the provider supports it | none |
| `response_format` | e.g. `{"type": "json_object"}` | none |

A value outside its range is **refused, not clamped**. Silently adjusting a number you chose would
hand you a call you did not ask for.

### What you may not pass, and why

You cannot set an `endpoint`, an `api_base`, an `api_version`, a credential reference, a `provider`,
or provider pass-through parameters. Wexa fills all of these from the model's own registry entry,
and supplying one is refused rather than ignored.

This is deliberate rather than incidental. The provider host is checked against your organization's
allow-list **at call time**, not when the model was registered, so withdrawing permission for a host
blocks a model that was registered while it was still permitted. Provider pass-through is
administered per model and checked before it reaches execution. A per-request version of either would
be an unchecked route into provider routing and credential placement. Your provider credential is
resolved at the point of execution and never reaches the caller.

### What comes back

```json
{
  "output": "pong",
  "model": "mdl_bx_anthropic-claude-haiku-4-5-20251001-v1_28ccc24bfca8",
  "source": "requested",
  "tokens": 28,
  "cost_usd": 0.000028,
  "latency_ms": 621
}
```

`model` is the model that **actually ran**, not the one you asked for, and `source` says why it was
chosen:

| `source` | Meaning |
| --- | --- |
| `requested` | You named this model |
| `organization default` | You named none, and this is the default the organization chose |
| `platform default` | You named none and the organization has chosen none, so a platform fallback ran |

Read both before attributing a charge. A `platform default` you did not expect is the signal that
nobody has chosen a default — the same empty setting described above, surfacing as a bill instead of
as a broken agent. On an enterprise deployment the call is refused outright in that state rather than
routed to a fallback, so no work runs on a model nobody selected.

### Errors, and which ones are worth retrying

Every failure carries a stable code, so your client can tell a mistake it made from a failure
upstream. The distinction that matters is terminal versus retryable: a client that cannot tell them
apart burns a full backoff cycle on a call that can never succeed.

| Code | Status | Meaning | Retry? |
| --- | --- | --- | --- |
| `S8:model-not-found` | 404 | No such model in this organization | No — fix the id |
| `S8:model-forbidden` | 403 | The model's host is no longer permitted | No — an administrator must act |
| `S8:model-invalid-parameters` | 422 | A parameter was outside its range or not settable | No — fix the request |
| `S8:model-call-limited` | 429 | This key's own model-call budget is full | Yes, shortly |
| `S8:model-rate-limited` | 429 | The **provider** rate-limited the call | Yes, with backoff |
| `S8:model-upstream-error` | 502 | The provider failed | Yes |

The two 429s are separate on purpose: one means you are asking Wexa for too much at once, the other
means your provider is. The provider's own message is passed through, because it is what makes a
credential or access failure diagnosable.

Both SDKs classify these for you and stop retrying the terminal ones.

### Limits and billing

Model calls are limited **separately** from ordinary API calls, and by how many are in flight rather
than by how many per minute — a long generation and a listing are not the same load. Exhausting your
model-call budget does not block your other calls, and exhausting the ordinary budget does not block
your model calls.

Every call is metered against the API key that made it, so per-key spend answers "which of my
integrations is spending this" before the invoice does. The spend draws down the same balance as
every other metered use; there is not a second one to watch.

### No streaming

`run-model` does not stream, and a request for it is refused with a clear message rather than served
the whole answer framed as a single chunk. A stream that does not stream invites interfaces built
against timing that would change the moment real streaming arrived, with no way to tell who depended
on the old behaviour.

## Platform-provided versus bring-your-own

Two ways to have a model at all, and the difference is entirely about billing and control.

**Platform-provided.** Wexa runs the model for you and meters the usage. There is no provider
account to set up and nothing to configure beyond choosing the model. Usage draws down **credits**,
the consumable balance your organization holds. The accounting is per call: an estimate is held
against your balance before the call is made, and reconciled afterwards against the tokens the call
actually used, for input and output separately. A call that fails after that hold gives the hold
back rather than keeping it. Running out of credits stops calls rather than letting them run up a
debt, so the failure is visible and bounded.

**Bring-your-own.** You supply your own provider credentials, and the provider bills you directly.
The cost is your existing account's cost, at your existing rates, and the rate limits that apply are
your account's rate limits rather than a platform-wide share. You see the spend where you already
see the rest of your provider spend.

### How the choice affects cost

- **Which model** you pick is the largest lever in both cases: an agent's bill is its tokens times
  the model's price, and models differ by more than the work usually does.
- **Platform-provided** turns cost into one number — credits — that is metered per call and drawn
  from a balance the organization tops up. One team's runs can exhaust the headroom another team was
  counting on, because the balance is organization-wide, in the same way a
  [quota](/docs/concepts/tenancy) is.
- **Bring-your-own** moves the cost off the credits balance entirely and onto your provider
  invoice. You trade the setup for direct visibility and for control of your own rate limits.
- **Where you set the model** decides the blast radius of a cost change. A per-agent `llm.model`
  moves one agent onto a more expensive model. `set-model` moves every agent that has not been
  overridden, in every project in the organization, at once.

## Where to go next

- [Agents](/docs/concepts/agents) — where `llm.model` lives.
- [Process flows](/docs/concepts/process-flows) — each agent in a manifest carries its own `llm`
  block, so one process flow can mix models deliberately.
- [Organizations, departments and projects](/docs/concepts/tenancy) — why the default is an
  organization-wide setting and why credits are counted there.
- [`run-model`](/docs/tools/run-model) — the full request contract for calling a model directly.
