---
title: "LLM routes"
description: "The OpenAI-compatible /v1 surface: which routes exist, how requests are translated and streamed, how fallbacks are chosen, and how upstream errors are mapped."
---
Point your existing OpenAI or Anthropic client at the gateway's base URL and
authenticate with an Obol virtual key. The gateway parses the request, decides
whether the model is allowed, picks the upstream, injects your own provider
credential, and passes the response through.
## Routes
Four routes accept LLM traffic. The path itself selects the input dialect.
| Path | Input format | Notes |
|---|---|---|
| `POST /v1/chat/completions` | OpenAI chat completions | The only path that can translate to a non-OpenAI provider |
| `POST /v1/responses` | OpenAI responses | Identity passthrough |
| `POST /v1/embeddings` | OpenAI embeddings | Identity passthrough |
| `POST /v1/messages` | Anthropic messages | Identity passthrough with model rewrite |
Any other `/v1` path is an `invalid_request` error.
`POST /v1/route` also lives under `/v1` but is not an LLM route. It is the
advisory capability resolver: it returns ranked capability cards and never
dispatches, unwraps a credential, or reaches a vendor. It authenticates under
the `mcp` scope, not the `llm` scope. See [Routing](/routing/overview).
Request bodies are limited to 8 MiB, and the whole request is bounded by a
timeout of `OBOL_UPSTREAM_TIMEOUT_S` plus 10 seconds, after which the gateway
returns `504`.
## Pipeline
Every LLM request runs the same ordered pipeline. Any step that fails renders an
error envelope in the caller's dialect and stops.
The `Authorization` header resolves to a principal, the workspace snapshot and
policy snapshot are pinned as one pair, and the request is rejected if the
workspace is frozen, the key's environment does not match the workspace's, or
the key lacks the `llm` scope. A `ob_test_` key never reaches a prod workspace.
The body is parsed into the input format's request type. A missing `model` is an
`invalid_request` error.
`model_visible` is the key's `allowed_models` glob list. A model that fails here
returns `model_not_allowed`, the same code an unroutable model returns — the
error does not distinguish "not allowed for you" from "not in this workspace".
`ModelRouter::resolve` prefers an exact snapshot entry, then the first matching
glob in snapshot order. The target carries the upstream model name, the
connection, a price entry, and any declared fallbacks.
The pinned policy publication is compiled (and cached) into an authorizer, and
the `complete` action is evaluated against the requested model. Anything but
`Allow` is `policy_denied`.
When the request streams, `stream_options.include_usage` is injected. An output
cap is injected when absent, and a checked token bound is reserved — for chat
completions the total output reservation is `n` times the per-choice cap, and
overflow is rejected before dispatch. The limiter then checks RPM and monthly
budget and reserves counters.
The upstream base URL passes the private-address guard, the connection must be
active and carry a credential, and `Vault::unwrap_checked` decrypts it (rejecting
an expired credential). The request is rendered for the target's wire format and
sent under a `UsageGuard`.
## Translation
The gateway supports a fixed set of translations. Anything else is rejected
rather than approximated.
| Input | Provider kind | Behavior |
|---|---|---|
| chat completions | `openai_compat` | Identity, model name rewritten to `upstream_model` |
| responses | `openai_responses` | Identity, model rewritten |
| embeddings | any | Identity; the upstream path comes from the input format |
| messages | `anthropic` | Identity, model rewritten |
| chat completions | `anthropic` | Translated to Anthropic Messages, and the response and stream translated back to OpenAI shapes |
The OpenAI-to-Anthropic translation defaults `max_tokens` to 4096, maps `user` to
`metadata.user_id` and `tool_choice: "required"` to `{"type": "any"}`, merges
consecutive same-role turns, and drops untranslatable top-level fields at debug
level. Any other combination is an unsupported-translation error.
## Streaming
A streaming response is passed through as `text/event-stream` with
`cache-control: no-cache`. For identity translations the SSE bytes are relayed
byte-exact; the gateway only *observes* frames to extract usage. SSE frames are
capped at 1 MiB so a hostile upstream cannot grow gateway memory.
Usage normalization across formats:
- `tokens_in` excludes cached tokens; `tokens_cached` is reported separately.
- `tokens_reasoning` is recorded when the provider reports it, and is never
billed twice.
- Anthropic's `message_delta.output_tokens` is cumulative and replaces rather
than accumulates.
The stream is marked complete on `[DONE]`, `response.completed`, or
`message_stop`. A cut stream drops the usage guard uncompleted, which emits the
usage event with `partial: true` and whatever was observed — a client disconnect
still bills what was consumed. For the translated chat-to-Anthropic stream,
`data: [DONE]` is emitted only after `message_stop`; an upstream error event, a
malformed event, or a premature EOF fails the stream with a safe error.
## Provider selection and fallbacks
Model resolution and provider selection happen against the workspace snapshot,
not at request time against a remote service. When the resolved target declares
fallbacks, they become a candidate list, and each candidate is checked
independently:
- It must pass CEL `model_visible` for this key.
- It must be `Allow` under Cedar `complete`, under the *same* pinned policy
publication as the primary.
- If the key has a monthly budget, every eligible candidate must be priced.
Unknown fallback pricing cannot silently become a zero-dollar reservation, so
the request fails `not_ready` instead.
The counter reservation is sized by the most expensive eligible candidate.
A fallback is attempted **only** on failures before the first byte: an upstream
`5xx`, a timeout, or a transport failure. A `4xx`, a parse failure, or anything
after the response has begun is returned to the caller as-is. Candidates whose
connection is inactive, missing, or missing a credential are skipped rather than
failing the request, unless they are the last candidate. If no candidate is
dispatched, the reservation is released.
See [Routing](/routing/overview) for how targets get into the snapshot.
## BYOK model keys
The gateway never holds a pooled provider account. Each model in the snapshot
names a connection, and the connection carries your encrypted provider
credential. Two properties follow:
- The plaintext key exists only between `Vault::unwrap_checked` and the outbound
socket. It is not serializable, its `Debug` prints `[REDACTED]` plus the last
four characters, and only that fragment is ever logged.
- The outbound request carries exactly `content-type` plus whatever the
credential scheme injects. Inbound headers are never forwarded, so a header
set by your client cannot reach your provider.
See [Vault](/security/vault) and [Virtual keys](/security/virtual-keys).
## Error mapping
Every failure renders one envelope. The `code` is stable across dialects; the
`type` field is translated into the shape the client expects.
| Code | HTTP | OpenAI `type` | Anthropic `type` |
|---|---|---|---|
| `unauthorized`, `key_revoked`, `key_expired` | 401 | `authentication_error` | `authentication_error` |
| `workspace_frozen`, `env_mismatch`, `model_not_allowed`, `tool_not_allowed`, `policy_denied` | 403 | `permission_error` | `permission_error` |
| `approval_required` | 423 | `permission_error` | `permission_error` |
| `budget_exceeded` | 402 | `insufficient_quota` | `permission_error` |
| `rate_limited` | 429 | `rate_limit_error` | `rate_limit_error` |
| `invalid_request` | 400 | `invalid_request_error` | `invalid_request_error` |
| `idempotency_conflict` | 409 | `invalid_request_error` | `invalid_request_error` |
| `no_eligible_route` | 422 | `invalid_request_error` | `invalid_request_error` |
| `upstream_error` | 502 | `server_error` | `api_error` |
| `upstream_timeout` | 504 | `server_error` | `api_error` |
| `not_ready` | 503 | `server_error` | `overloaded_error` |
| `internal` | 500 | `server_error` | `api_error` |
Every error body also carries `request_id` and `retryable`, plus `retry_after_s`
and `vendor_status` when known. Rate-limited responses set the `retry-after`
header. Every response — success or failure — carries `x-request-id`.
### What the gateway will not repeat back
Upstream error bodies are used for mapping, not for relaying:
- On `401`/`403` and on `5xx`, the envelope carries the fixed message
`upstream rejected credentials` (or a server-error equivalent) plus
`vendor_status`, never the vendor body — some providers echo the key.
- On `400`, the upstream `error.message` is surfaced, because it describes the
caller's own request.
- Transport-error messages are classification-only, because the underlying error
display includes the URL, which can carry a query-scheme secret.