Models
A model is the LLM an agent uses: either one the platform provides or one you bring. Every agent needs one, and the choice has two consequences — what the agent is capable of, and who pays for it.
How an agent gets its model
An agent's definition carries an llm block whose model names the LLM. It defaults to the
sentinel value system_model, which does not name a model at all: it means whichever model this
organization has chosen.
That gives two levers, at two different levels.
- Per agent. Put a real model id in that agent's
llm.model. Only that agent is affected, and it stops following the organization default. - Per organization.
set-modelchanges the default itself. Every agent still usingsystem_modelfollows it immediately, across every project in the organization. Agents that name their own model are unaffected.
Because the second reaches every project, it requires an admin role. It also takes an optional
fallback — a second model used when the default is unavailable.
Finding out which models you can use
list-models takes no arguments — which organization is being asked
about comes from your credential — and returns the models this organization can run, along with
which one is the default.
It answers with two families of id, and both are valid in an agent's llm.model.
- The registry entries are the organization's own model registry: an id, a name, a provider, a health status, and whether it is the default. These are the ones to prefer.
- The available list is what the running service reports it can reach: the
system_modelsentinel, plus composite ids identifying a particular model from a particular provider for a particular organization.
The registry is read from a separate service, so it is possible for list-models to answer while
that part is unavailable. When that happens the result says so rather than returning a short list
that looks complete. Treat a missing registry as "not known right now", not as "this organization
has no registry models".
set-model refuses a model id the organization does not actually have configured, and refuses an
unconfigured fallback too. That refusal is doing you a favour: an accepted id that resolves to
nothing would break every agent in the organization on its next run, and a rejected call is a much
cheaper way to find out.
Calling a model directly
Everything above is about which model an agent uses. You can also call a model yourself, with no
agent involved: run-model takes a model and a prompt and returns the
answer, the tokens it consumed, what it cost and how long it took.
What you pass
model accepts any id list-models returns — a registry id or a composite — and is optional.
Omit it and the call runs the organization's default, resolved exactly as it is for an agent on
system_model. A free-text provider name such as gpt-4o is refused: a model has to be one this
organization has registered, so that every call resolves to a real credential and a permitted host.
Send either a prompt — the simple form, treated as one user message — or a messages list of
role/content pairs for a system message and a multi-turn conversation. When both are present,
messages wins. The roles are system, user and assistant; there is no tool role, because tool
calling is not offered here.
These are the controls you may set, and they are an allow-list rather than a pass-through:
| Field | Meaning | Default |
|---|---|---|
max_tokens | Ceiling on the tokens generated | 512 |
temperature | 0 to 2 | 0.2 |
top_p | Nucleus sampling, 0 to 1 | provider's |
stop | Up to four sequences that end generation | none |
seed | Best-effort determinism, where the provider supports it | none |
response_format | e.g. {"type": "json_object"} | none |
A value outside its range is refused, not clamped. Silently adjusting a number you chose would hand you a call you did not ask for.
What you may not pass, and why
You cannot set an endpoint, an api_base, an api_version, a credential reference, a provider,
or provider pass-through parameters. Wexa fills all of these from the model's own registry entry,
and supplying one is refused rather than ignored.
This is deliberate rather than incidental. The provider host is checked against your organization's allow-list at call time, not when the model was registered, so withdrawing permission for a host blocks a model that was registered while it was still permitted. Provider pass-through is administered per model and checked before it reaches execution. A per-request version of either would be an unchecked route into provider routing and credential placement. Your provider credential is resolved at the point of execution and never reaches the caller.
What comes back
{
"output": "pong",
"model": "mdl_bx_anthropic-claude-haiku-4-5-20251001-v1_28ccc24bfca8",
"source": "requested",
"tokens": 28,
"cost_usd": 0.000028,
"latency_ms": 621
}
model is the model that actually ran, not the one you asked for, and source says why it was
chosen:
source | Meaning |
|---|---|
requested | You named this model |
organization default | You named none, and this is the default the organization chose |
platform default | You named none and the organization has chosen none, so a platform fallback ran |
Read both before attributing a charge. A platform default you did not expect is the signal that
nobody has chosen a default — the same empty setting described above, surfacing as a bill instead of
as a broken agent. On an enterprise deployment the call is refused outright in that state rather than
routed to a fallback, so no work runs on a model nobody selected.
Errors, and which ones are worth retrying
Every failure carries a stable code, so your client can tell a mistake it made from a failure upstream. The distinction that matters is terminal versus retryable: a client that cannot tell them apart burns a full backoff cycle on a call that can never succeed.
| Code | Status | Meaning | Retry? |
|---|---|---|---|
S8:model-not-found | 404 | No such model in this organization | No — fix the id |
S8:model-forbidden | 403 | The model's host is no longer permitted | No — an administrator must act |
S8:model-invalid-parameters | 422 | A parameter was outside its range or not settable | No — fix the request |
S8:model-call-limited | 429 | This key's own model-call budget is full | Yes, shortly |
S8:model-rate-limited | 429 | The provider rate-limited the call | Yes, with backoff |
S8:model-upstream-error | 502 | The provider failed | Yes |
The two 429s are separate on purpose: one means you are asking Wexa for too much at once, the other means your provider is. The provider's own message is passed through, because it is what makes a credential or access failure diagnosable.
Both SDKs classify these for you and stop retrying the terminal ones.
Limits and billing
Model calls are limited separately from ordinary API calls, and by how many are in flight rather than by how many per minute — a long generation and a listing are not the same load. Exhausting your model-call budget does not block your other calls, and exhausting the ordinary budget does not block your model calls.
Every call is metered against the API key that made it, so per-key spend answers "which of my integrations is spending this" before the invoice does. The spend draws down the same balance as every other metered use; there is not a second one to watch.
No streaming
run-model does not stream, and a request for it is refused with a clear message rather than served
the whole answer framed as a single chunk. A stream that does not stream invites interfaces built
against timing that would change the moment real streaming arrived, with no way to tell who depended
on the old behaviour.
Platform-provided versus bring-your-own
Two ways to have a model at all, and the difference is entirely about billing and control.
Platform-provided. Wexa runs the model for you and meters the usage. There is no provider account to set up and nothing to configure beyond choosing the model. Usage draws down credits, the consumable balance your organization holds. The accounting is per call: an estimate is held against your balance before the call is made, and reconciled afterwards against the tokens the call actually used, for input and output separately. A call that fails after that hold gives the hold back rather than keeping it. Running out of credits stops calls rather than letting them run up a debt, so the failure is visible and bounded.
Bring-your-own. You supply your own provider credentials, and the provider bills you directly. The cost is your existing account's cost, at your existing rates, and the rate limits that apply are your account's rate limits rather than a platform-wide share. You see the spend where you already see the rest of your provider spend.
How the choice affects cost
- Which model you pick is the largest lever in both cases: an agent's bill is its tokens times the model's price, and models differ by more than the work usually does.
- Platform-provided turns cost into one number — credits — that is metered per call and drawn from a balance the organization tops up. One team's runs can exhaust the headroom another team was counting on, because the balance is organization-wide, in the same way a quota is.
- Bring-your-own moves the cost off the credits balance entirely and onto your provider invoice. You trade the setup for direct visibility and for control of your own rate limits.
- Where you set the model decides the blast radius of a cost change. A per-agent
llm.modelmoves one agent onto a more expensive model.set-modelmoves every agent that has not been overridden, in every project in the organization, at once.
Where to go next
- Agents — where
llm.modellives. - Process flows — each agent in a manifest carries its own
llmblock, so one process flow can mix models deliberately. - Organizations, departments and projects — why the default is an organization-wide setting and why credits are counted there.
run-model— the full request contract for calling a model directly.