Routing
Learn how LLMGateway intelligently routes your requests to the best available models and providers.
LLMGateway provides flexible and intelligent routing options to help you get the best performance and cost efficiency from your AI applications. Whether you want to use specific models, providers, or let our system automatically optimize your requests, we've got you covered.
LLMGateway also includes automatic retry and fallback — if a provider fails, your request is seamlessly retried on the next best provider, all within the same API call.
Routing Analytics
Admin routing analytics includes static catalogue mappings and Airside listings, including mappings with no traffic in the selected window. Airside listings use their filed prices. Inactive listings remain visible but are excluded from scoring. Approved, active Airside carriers participate in automatic provider selection and fallback using their filed prices and capabilities, including carriers defined entirely in Airside. Inactive or unapproved carriers remain visible but are excluded from scoring. A provider prefix is optional, including when multiple carriers list a model that is absent from the static catalogue.
Airside carriers follow the same credential, capability, policy, rate-limit, and
session rules as other providers. Adding a carrier does not expand the models
that auto may choose. Dynamic-route provider lists include active Airside
carriers; availability is checked when saving, publishing, rolling back, and
executing a route.
Scores use the default routing config. They are computed on the router's live
metrics window and on hourly averages, for each request shape that changes the
weights: streaming, cached prompts, sessions, and routing preferences.
Global Provider Rate Limits
Administrators can set global or organization-specific RPM/RPD caps. Global limits support Global (shared) or Per-organization counters. Changes propagate through the rate-limit cache (default 60 seconds).
Organization-specific limits and more specific global provider/model limits can override a provider-wide limit for the same window. No configured limit means unlimited.
Each limit is strict (the default), soft, or lax. Soft and lax
limits allow sessions already pinned to the capped
provider to keep using it. Lax also allows requests that explicitly name that
provider (provider/model), even without a session or on its first request.
| Request after reaching the cap | Soft | Lax |
|---|---|---|
| Existing session pinned to the provider/model | Allowed | Allowed |
| Explicit provider request without an existing session pin | No exemption | Allowed |
| Automatic routing without an existing session pin | Routed away | Routed away |
| Fallback from another provider | No exemption | Excluded |
Exempt requests still count toward the limit. A session id alone does not grant an exemption, and every exceeded RPM/RPD window must permit the request: a strict window grants no exemption, while a soft window still requires an existing session pin.
A limit of 0 permits only requests exempt under its mode: strict blocks
all matching requests, soft allows existing session pins, and lax also
allows explicit provider requests. Zero caps cannot be used as a last resort
for non-exempt traffic. Routing can choose an available alternative; if no
eligible initial route remains, the gateway returns 429. Exempt requests
still increment both configured counters.
Lax caps apply to automatic provider selection under every routing strategy,
including dynamic provider ordering. Auto model selection excludes exhausted
routes before ranking models, so another eligible model can serve the request.
A capped lax provider cannot be used as a
last resort when all candidates are capped; requests with no eligible initial
route return 429. Retries on the originally selected provider retain its
exemption, including credential and region changes; cross-provider fallback
does not. Administrators can configure lax limits globally or for a specific
organization, with the same precedence and counter scopes as other modes.
Model Selection
Any Model Name
You can use any model name from our models page or discover available models programmatically through the /v1/models endpoint.
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}]
}'Model ID Routing
Choose a specific model ID to route to the best available provider for that model. LLMGateway's smart routing algorithm considers multiple factors to find the optimal provider across all configured options.
Smart Routing Algorithm
When you use a model ID without a provider prefix, LLMGateway's intelligent routing system analyzes multiple factors to select the best provider.
Weighted Scoring System:
Each factor has a relative weight. The factors are scored as ratios against the best provider in the candidate set (e.g. a provider that is twice as expensive as the cheapest scores 1.0 on price), and each ratio is multiplied by its weight divided by the sum of all active weights. The provider with the lowest (best) total score wins.
The default weights are:
| Factor | Default weight | Notes |
|---|---|---|
| Price | 0.6 | Token cost for the expected input/output mix, including cache reads when relevant |
| Uptime | 0.5 | Provider reliability / low error rate |
| Throughput | 0.05 | Tokens per second generation speed |
| Latency | 0.025 | Time to first token — only applied for streaming requests |
| Cache | 0 | Optional cache-support preference; cache-read savings already count toward price |
| Image price | 1.0 | Replaces the price weight for image-generation models |
Because the weights are relative and normalized by the sum of the active weights, price and uptime dominate routing decisions in practice, while throughput and latency act as tie-breakers between otherwise comparable providers.
Latency Weight for Non-Streaming Requests:
The latency weight only applies to streaming requests (time-to-first-token is only measured there). For non-streaming requests the latency weight is dropped and its share is redistributed proportionally across the remaining factors.
Time-Decayed Metrics Window:
Provider metrics (uptime, throughput, latency) are not a flat "last N minutes" snapshot. They are aggregated over a rolling 60-minute window with a time-decay weighting so very recent behavior dominates while older data still contributes:
- The most recent 1 minute is weighted 10×
- The most recent 5 minutes are weighted 3×
- The remainder of the 60-minute window is weighted 1×
This makes routing react quickly to a provider that just started failing or slowing down, without overreacting to a single noisy data point.
Only requests served on LLM Gateway's own credentials feed these metrics. Requests on your own provider keys never count toward a provider's uptime, so a misconfigured key cannot degrade routing for anyone else.
Prompt Caching and Token Costs:
For estimated prompts of at least 5,000 tokens, or when choosing a session's provider, routing blends each provider's uncached and cached input prices and weights output by the expected output:input token ratio. Coding sessions that mostly reuse prompt tokens can therefore favor a provider with cheaper cache reads even when its uncached input price is higher. Providers without a cached input price use their full input price.
These estimates use the project's last 24 hours of model usage, once it includes at least 20 successful requests and 20,000 input tokens:
- Cache-hit rate is cached input tokens divided by total input tokens. Each provider uses its own rate, counted across all of its regions, once it meets the same sample thresholds; otherwise it uses the project's combined rate for that model.
- Output:input ratio uses the project's combined output and input tokens for that model across providers.
Routing caches the usage lookup for 60 seconds. It reads hourly aggregates, which work with payload retention disabled. Hourly buckets containing gateway response-cache hits are excluded because they cannot isolate upstream usage, so projects with response caching enabled may keep using the defaults. During lookup failures, routing uses previously cached observations when available, then falls back to configured estimates. These are workload estimates, not guarantees that a particular prompt will hit a provider's cache.
Without enough history, routing uses these initial workload estimates:
| Workload | Cached input | Output:input ratio |
|---|---|---|
| General API / unknown | 10% | 20% |
| DevPass or a recognized coding client | 90% | 2% |
| Chat organization | 50% | 10% |
Recognized coding clients use the coding defaults even on regular API projects. A session id alone does not identify coding traffic. These are starting assumptions; sufficient project/model observations replace them. The dashboard reports organization defaults, while recognized coding requests use the coding profile at request time.
Explicit Enterprise overrides for thresholds.cacheHitRate and thresholds.cacheOutputRatio take precedence over both workload defaults and observations. Setting them to 0 and 1, respectively, restores list-price ranking.
Both auto and price routing use these token-cost estimates. The separate cache weight defaults to 0, so cache support alone does not outweigh lower estimated costs. Enterprise projects can explicitly enable that additional preference under auto; price routing always sets it to zero. Cache support appears as cacheSupported in routing metadata.
When choosing a session's provider, routing applies the workload estimate even to a short opening prompt, using observations when available and workload defaults otherwise. This estimates the session's token mix; the opening request may still incur cache misses. Small requests outside a session weight input and output prices equally and omit the cache weight.
Exponential Uptime Penalty:
Providers with uptime below 95% receive an additional exponential penalty that increases rapidly as uptime drops:
- 95-100% uptime: No penalty
- 90% uptime: ~0.07 penalty
- 80% uptime: ~0.62 penalty
- 70% uptime: ~1.73 penalty
- 50% uptime: ~5.61 penalty
This ensures providers experiencing significant issues are strongly deprioritized while minor fluctuations have minimal impact. The penalty threshold (default 95%) is configurable.
Provider Priority:
Each provider has a priority value that nudges routing toward or away from it independently of live metrics. Several providers ship with their own built-in priority in the catalogue; a provider without one defaults to 1, and Enterprise projects can override the value per provider (see Per-Project Routing Configuration):
- A provider's priority is applied as a
(1 - priority)adjustment to its score — higher priority lowers the score (more preferred), lower priority raises it (less preferred). - A priority of 0 disables the provider entirely, removing it from routing for that model.
Provider priorities are surfaced in the routing metadata so you can see how they influenced a decision.
Epsilon-Greedy Exploration (1% of requests by default):
To solve the "cold start problem" where new or unused providers never get traffic to build up metrics, the system randomly explores different providers a small fraction of the time (default 1%, configurable). This ensures:
- All providers periodically receive traffic
- New providers can prove their reliability
- The system adapts to changing provider performance
- You benefit from improved routing decisions over time
The exploration rate is configurable per project through the routing configuration (thresholds.explorationRate), and self-hosted deployments can override it globally with the EXPLORATION_RATE environment variable (a number between 0 and 1).
Models with provider-bound reasoning never explore. OpenAI encrypted reasoning and Gemini thought signatures can only be verified by the provider that issued them, so a turn randomly sent elsewhere would fail with an error such as invalid_encrypted_content.
Stable Provider Preference:
To avoid unnecessary churn between providers that score similarly, LLMGateway remembers the best provider chosen for each model and sticks with it across requests — even if another provider edges ahead slightly on the next score calculation.
On every routing decision, the system checks whether the previously selected provider is still acceptable:
- Uptime hard switch: if the preferred provider's uptime drops below 85%, routing switches to the current best-scoring provider immediately.
- Score margin soft switch: the preferred provider is replaced only when a better option's score is more than 0.15 ahead. Small fluctuations caused by metric noise or minor price differences do not trigger a switch.
- Periodic re-evaluation: the preference expires after 1 hour, at which point the next request picks the best-scoring provider fresh and stores it as the new preferred.
Requests that are part of the epsilon-greedy exploration bypass this preference entirely so that all providers continue to receive periodic traffic and build up metrics.
The selection reason in routing metadata will show stable-preferred when a request was served by the stored preference rather than the top-scored provider at that moment.
The preference lifetime (default 3600 seconds), hard-switch uptime floor
(default 85), and soft-switch score gap (default 0.15) come from the
routing configuration's sticky settings. On the Enterprise plan, they
can be customized per project from the dashboard — see Per-Project Routing
Configuration.
Your Own Provider Keys (hybrid mode):
In a project in hybrid mode, routing first narrows the candidates to providers you have your own provider key for, and the weighted score picks among those. LLM Gateway's own credentials are used when none of your keys can serve the requested model, as fallback when attempts on your keys fail, or when your current keyed provider's uptime is degraded and no keyed candidate is healthier — escaping the degraded provider takes priority over staying on your key. Projects in API-keys or credits mode are unaffected.
Dynamic routes can specify provider preference order.
After eligibility checks and hybrid-mode key preference, that order controls the
initial provider and cross-provider fallback on non-sticky requests. A healthy
session pin still wins; if it becomes unavailable, the next request selects a
replacement using the order. Sticky requests still skip cross-provider retries.
Regional selection within a provider still uses scores. These requests skip
random exploration and record provider-order when the order decides selection.
Routing Metadata:
Every request includes detailed routing metadata in the logs, showing:
- Available providers that were considered
- Selected provider and selection reason
- Scores for each provider (including uptime, throughput, latency, price, priority, and cache support)
This transparency allows you to understand and debug routing decisions.
Using model IDs without a provider prefix automatically routes to the optimal provider based on reliability, speed, and cost. The system continuously learns and adapts based on real-time performance metrics.
Smart routing prioritizes reliability over cost, ensuring your requests are routed to providers with proven uptime and performance, while still considering cost efficiency.
Routing Strategy
By default, model-ID routing uses the full weighted score described above (routing: "auto"). When you care about a single dimension, set the routing field — named after the factor it optimizes — to bias provider selection toward it. Each project can also set a default strategy under Settings → Routing (on coding/dev plans, only auto and price can be set); a per-request routing value overrides the project default:
| Strategy | Behavior |
|---|---|
auto (default) | Full weighted smart-routing score (price, uptime, throughput, latency, cache). |
price | Gives price a 90% relative weight, including estimated cache-read costs when relevant. |
throughput | Gives throughput a 90% relative weight, so the fastest-generating provider wins. |
latency | Gives latency a 90% relative weight, so the lowest time-to-first-token wins. |
Each non-auto strategy keeps a small (10%) uptime weight, and the exponential uptime penalty still applies on top. This means the dominant pick is still skipped in favor of another provider when it has extremely bad uptime — you get the cheapest (or fastest) provider that is actually healthy, not one that is effectively down.
A non-auto strategy also turns off epsilon-greedy exploration for the request: you asked to optimize a single factor, so no traffic is randomly rerouted to other providers. (A global EXPLORATION_RATE environment override on self-hosted deployments still takes precedence.)
Because time-to-first-token is only measured for streaming requests, routing: "latency" only biases streaming requests; for non-streaming requests it falls back to selecting on uptime.
# Always pick the cheapest healthy provider for this model
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v3.2",
"messages": [{"role": "user", "content": "Hello!"}],
"routing": "price"
}'# Always pick the highest-throughput healthy provider for this model
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v3.2",
"messages": [{"role": "user", "content": "Hello!"}],
"routing": "throughput"
}'The routing field only applies to model-id routing. Combining it with a
specific provider (e.g. openai/gpt-4o) returns a 400 error, since the
strategy can't influence a pinned provider — remove the provider prefix to use
a strategy. On coding (dev) plans, only auto and price are allowed;
the other strategies return a 400 error because they would bypass the
prompt-cache–aware routing those plans depend on.
Anthropic safeguards
A nonempty safeguards body with its paired dated beta makes server-side review
support a routing requirement. Automatic model and provider selection, session
pins, and retries exclude providers that cannot forward it. An explicit
incompatible provider pin, or automatic selection with no compatible provider,
returns 400, including for streaming requests. Routing preferences cannot
relax this requirement; a beta header alone does not enable it.
See Claude Code auto mode for upstream availability and the opt-out that restores broader routing with billed client-side classifier requests.
Sticky Session Routing
When a model is served by multiple providers, every request is normally scored independently — so a multi-turn conversation can bounce between providers. That defeats provider-side prompt caching, which only pays off when consecutive requests with a shared prefix hit the same provider.
Sticky session routing solves this: attach a session identifier and LLMGateway pins all requests for that session to a single provider (and region), keeping the upstream prompt cache warm across the whole conversation.
Setting the session id
For chat completions, the session key is resolved in priority order:
- The
x-session-idheader - The
x-session-affinityheader (sent automatically by coding agents such as opencode) - The
session_idorsession-idheader - The
prompt_cache_keybody field (OpenAI-compatible) - The
userbody field (OpenAI-compatible)
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-H "x-session-id: conversation-9f8e7d6c" \
-d '{
"model": "claude-sonnet-4-6",
"messages": [{"role": "user", "content": "Hello!"}]
}'For the Anthropic Messages endpoint (/v1/messages), the session key is derived automatically from metadata.user_id — coding agents such as Claude Code embed the session id there — and forwarded internally. An explicit x-session-id header still takes precedence.
How pinning works
On a session's first request, a dynamic route with explicit provider order chooses the first eligible provider and scores its regions. Otherwise, the normal weighted smart-routing score chooses the provider using price, priority, uptime, and throughput. That choice is then persisted for the session and reused on every subsequent request, so the upstream prompt cache stays warm without bouncing the conversation between providers.
The first selection uses the expected cache-hit rate and output/input mix even if the opening prompt is short. Routing uses observed project/model usage when sufficiently sampled, otherwise the workload defaults above. New usage observations affect future provider selections; a healthy existing pin stays in place as described below.
Because the pinned provider is replayed directly, sticky requests skip the epsilon-greedy exploration — a session is never randomly bounced to a different provider mid-conversation.
Request compatibility takes precedence over the saved pin. The gateway first filters mappings for requirements such as input modalities, service tiers, regions, requested Anthropic safeguards, and a non-auto tool_choice, then looks for the pinned provider in that eligible set. If the pinned mapping cannot honor the request but another mapping can, the session moves to the capable mapping and the pin is updated. For a fixed model or dynamic route where no mapping can honor tool_choice, the gateway preserves availability instead: it keeps the mappings, downgrades tool_choice to auto, and sticky routing may retain the existing pin. Automatic model selection does not use that fallback because it can choose a capable model instead.
Falling back when a provider is down
An established pin yields only when its provider can no longer serve the session well. A session selects and pins a replacement when its provider:
- Drops below the session uptime threshold (default 85%), except for sessions on models with provider-bound reasoning, or
- Is filtered out of the candidate set (health or compatibility filtering).
Sticky requests never enter the cross-provider automatic retry & fallback loop — a transient failure is retried against the pinned provider only, on another configured key when several exist, or on the same platform key when only one is configured. The failure still degrades that provider's uptime metrics, which is what triggers re-pinning on a subsequent request once the uptime threshold is crossed.
Re-pinning follows the same selection rules as the first request: explicit dynamic-route order chooses the provider and scores choose its region; otherwise weighted scoring chooses the provider.
Sessions on models with provider-bound reasoning keep their eligible provider through uptime dips, because OpenAI encrypted reasoning and Gemini thought signatures cannot be replayed on another provider. If a pin expires or the provider becomes ineligible, moving a conversation with its old signatures can produce Corrupted thought signature or Invalid thought signature. On the pay-as-you-go API, keep a Gemini conversation on one provider with a provider-prefixed model and X-No-Fallback: true from its first request. For an affected conversation, resume with the provider that issued its signatures or start a new conversation. See reasoning replay.
The selection reason in routing metadata shows session-sticky when a request was pinned via a session id.
Sticky routing optimizes for cache locality over per-request churn. Once a session is pinned it stays on its provider even if a cheaper or faster alternative becomes momentarily available, since the prompt-cache savings typically outweigh the difference — but the initial pick still respects price and priority. Requests without a session id are unaffected and continue to use the weighted smart-routing algorithm.
Provider-Specific Routing
To use a specific provider without any fallbacks, prefix the model name with the provider name followed by a slash:
# Use OpenAI specifically
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}]
}'
# Use DeepSeek provider specifically
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v3.2",
"messages": [{"role": "user", "content": "Hello!"}]
}'Provider-specific routing is not available on DevPass coding plans —
provider-prefixed model ids (and custom provider routing) return a 403
there. DevPass always uses the prompt-cache–aware smart routing above with
plain model ids. Provider pinning requires the pay-as-you-go API. See
Provider routing on
DevPass.
Regions
Some providers expose the same model in multiple regions. In that case, LLMGateway supports two routing modes:
provider/modelselects the best eligible region for that provider using the same routing inputs used elsewhere: recent uptime, throughput, latency, and priceprovider/model:regionpins the request to one exact region
When Airside owns a provider/model pair, its active regional listings replace the catalogue regions. A pin to an unlisted, inactive, or expired region is rejected. Scheduled regional deactivations apply even while listings are cached.
# Let LLMGateway choose the best Alibaba region for DeepSeek V3.2
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "alibaba/deepseek-v3.2",
"messages": [{"role": "user", "content": "Hello!"}]
}'
# Force a specific Alibaba region
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "alibaba/deepseek-v3.2:cn-beijing",
"messages": [{"role": "user", "content": "Hello!"}]
}'If your provider key stores an explicit region, that region acts like a lock and LLMGateway will only use that region for provider-specific requests. If no explicit region is configured on the provider key, provider-specific requests can still score all eligible regions for that provider.
Routing metadata reflects this:
- Dynamic provider-region selection shows all eligible regional scores that were considered
- Explicitly pinned regions show only the pinned region in the score list
Region-aware routing only compares regions that are actually available for the current project mode and provider setup. In credits mode, that means only regions backed by configured environment keys. In API keys and hybrid mode, an explicit provider-key region restricts the request to that region.
A few regions are served by an endpoint that belongs to your own account rather than a shared one — Alibaba Cloud's EU (Frankfurt) region has no shared DashScope domain and is served by your Model Studio workspace's dedicated host. Such a region still works from an API key alone, via the provider's shared entry point, but that endpoint is rate-limited and carries no SLA. Set the workspace ID on the provider key (copy it from the API Host shown when you create the key) to route through your own endpoint instead.
Low-Uptime Protection
When you specify a provider explicitly, LLMGateway checks the provider's recent uptime (from the time-decayed metrics window described above). If the uptime falls below 90%, the system automatically routes your request to the best available alternative provider to ensure reliability. This protects your application from providers experiencing temporary issues. The fallback threshold (default 90%) is configurable. Models with provider-bound reasoning skip this reroute, since the alternative provider would reject the conversation's reasoning payloads.
If the requested provider has low uptime but no alternative providers are available for that model, the request will still be sent to the originally requested provider.
Disabling Fallback with X-No-Fallback Header
If you need to bypass this protection and always use the exact provider you specified regardless of its current uptime, you can use the X-No-Fallback header:
# Force use of a specific provider even if it has low uptime
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-H "X-No-Fallback: true" \
-d '{
"model": "openai/gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}]
}'Using X-No-Fallback: true disables automatic provider failover. Your
requests will be sent to the specified provider even if it is experiencing
issues, which may result in higher error rates. Retries may still occur
against another key for the same provider when multiple keys are configured.
When the X-No-Fallback header is used, the routing metadata in logs will include noFallback: true to indicate that fallback was disabled for that request.
Automatic Retry & Fallback
When using model ID routing (without a provider prefix), LLMGateway automatically retries failed requests on alternate providers. This happens transparently within the same API call — your application receives the successful response as if nothing went wrong.
How Retry Works
- Your request is routed to the best available provider using the smart routing algorithm
- If that provider fails with a retryable error (see What Triggers a Retry below), the gateway marks the provider as failed
- The next best available provider is selected and the request is retried
- Up to 2 retries by default (configurable per project via the routing configuration) are attempted before returning an error to the client
Request → Provider A (500 error) → Provider B (200 OK) → ResponseBoth streaming and non-streaming requests support automatic retry. A streaming request is retried when the provider fails before sending any output, including an error event that follows only stream lifecycle events such as response.created. Once output has reached the client, a later error ends the stream instead.
What Triggers a Retry
Retries are triggered by failures classified as provider-side or gateway-side problems:
- 5xx errors (500 Internal Server Error, 502 Bad Gateway, 503 Service Unavailable, etc.)
- Timeouts (upstream provider took too long to respond)
- Connection failures (network errors, DNS failures, etc.)
- Upstream rate limits (a
429from the provider) - Provider account and mapping problems — an upstream
401/403(bad provider credentials),402(provider account out of funds),404/405(model or endpoint mapping gap), and a few specific400bodies that indicate the same kinds of gateway-side problems
Retries are not triggered by:
- 4xx client errors — a request that is genuinely invalid (validation errors, unsupported parameters) fails the same way everywhere, so it is passed through with its original status instead of retried
- Content filter responses (Azure ResponsibleAI, etc.)
When Retry Is Disabled
Automatic retry to a different provider is disabled when:
- The
X-No-Fallback: trueheader is set - A specific provider is requested (e.g.,
openai/gpt-4o) - The request carries a session id and sticky session routing is enabled — the session stays pinned to its provider
- The request replays provider-bound reasoning, such as encrypted reasoning or signed thought parts — another provider would reject those payloads. This pin applies regardless of the selected mapping's reasoning capability. A reasoning-capable model alone does not disable retries: requests without replayed payloads can fail over before output starts.
- No alternative providers are available for the requested model
- The maximum retry count (default 2) has been exhausted
Retries can still happen within the same provider when multiple keys are configured and the current key fails with a retryable error.
Routing Transparency
Every provider attempt — both failed and successful — is recorded in the routing array in the response metadata (streaming and non-streaming alike) and activity logs:
{
"metadata": {
"routing": [
{
"provider": "openai",
"model": "gpt-4o",
"status_code": 500,
"error_type": "server_error",
"succeeded": false,
"credentialSource": "byok",
"apiKeyHash": "f029ee9",
"providerKeyId": "pk_2f9a...",
"providerKeyLabel": "billing-team-key"
},
{
"provider": "azure",
"model": "gpt-4o",
"status_code": 200,
"error_type": "none",
"succeeded": true,
"credentialSource": "platform",
"apiKeyHash": "ecb88d5"
}
]
}
}Whose key served each attempt
credentialSource says who owns the provider credential an attempt was sent with:
| Value | Meaning |
|---|---|
byok | Your own provider key. The provider bills you directly and the attempt is not deducted from your credits. |
platform | An LLM Gateway credential. The attempt runs on credits and is deducted from your balance. |
This matters most in hybrid mode, where a request that fails on your own key falls back to LLM Gateway's credential: both attempts appear in the same routing array, and only credentialSource tells them apart — apiKeyHash is an opaque fingerprint that says two attempts used different keys, not which key was yours. The same value is stored on the log as routingMetadata.usedCredentialSource for the credential that ultimately served the request, and is shown as a your key / LLM Gateway key badge in the dashboard's routing view.
Which of your keys ran
A byok attempt also carries the key itself: providerKeyId, and providerKeyLabel — the key as it is named on your provider keys page (its name, or its masked token when it has none). So when several of your keys are configured for a provider and the gateway rotates between them, each attempt says which one it used instead of leaving you to decode a fingerprint.
Chat requests additionally record routingMetadata.eligibleProviderKeys on the log: your keys that were candidates for the provider that served the request, in selection order. It is omitted for credits-mode projects, which route on LLM Gateway credentials, and for custom providers, whose keys are scoped by their own catalogue.
These fields describe your keys only. LLM Gateway's own credentials — the
ones that serve credits-mode traffic — are never named: a platform attempt
still reports credentialSource and apiKeyHash, but never providerKeyId
or providerKeyLabel.
Retried Log Tracking
Each provider attempt creates its own log entry. Failed attempts that were retried are marked with:
retried: true— indicates this failed request was retried on another providerretriedByLogId— the ID of the final successful log entry
This allows you to distinguish between unrecovered failures and failures that were transparently recovered via retry. In the dashboard, retried logs display a "Retried" badge with a link to the successful log.
Impact on Provider Health
Failed attempts still count against the provider's uptime score, even when the request was successfully retried on another provider. This means:
- A provider that keeps failing will see its uptime score drop
- Only gateway and upstream errors count: requests rejected as client errors (invalid request bodies, unsupported parameters) are excluded from both the error count and the request total, so your own bad requests never mark a provider as down
- The exponential uptime penalty kicks in below 95% (see Smart Routing Algorithm)
- Future requests are automatically routed away from unreliable providers
- A custom provider routed by bare model id is skipped while its uptime is below the low-uptime threshold and another provider is available
- Your application stays reliable without any code changes on your side
Automatic retry and fallback works together with smart routing to provide self-healing behavior. Failing providers are automatically avoided, and your requests are transparently recovered on reliable alternatives.
Per-Project Routing Configuration (Enterprise)
All plans use observed token usage for cache pricing when sufficient history exists. On the Enterprise plan, you can override the settings listed below per project from the dashboard under Project Settings → Routing, including explicit cache-pricing assumptions. The 24-hour usage window and minimum sample requirements are fixed; the History settings control uptime, throughput, and latency metrics.
Overrides are merged on top of the defaults, so you only set the values you want to change. When a custom configuration is disabled, the project falls back to the defaults.
The following groups can be customized per project:
| Group | What it controls | Defaults |
|---|---|---|
| Weights | Relative importance of each scoring factor | price 0.6, imagePrice 1.0, uptime 0.5, throughput 0.05, latency 0.025, cache 0 |
| Thresholds | Cache prompt size and pricing overrides, uptime-penalty threshold, exploration rate, and fallback metrics | cachePromptTokens 5000, cacheHitRate 0.1, cacheOutputRatio 0.2 (coding: 0.9 / 0.02; Chat: 0.5 / 0.1), uptimePenalty 95, defaultUptime 100, defaultLatency 1000, defaultThroughput 50, explorationRate 0.01 |
| Retry | Max cross-provider fallback attempts and the low-uptime reroute threshold | maxRetries 2, lowUptimeFallbackThreshold 90 |
| Timeouts | Per-request time limits (streaming, non-streaming, and the gateway value that caps the default streaming limit) — see Request Timeouts. Capped at the infrastructure defaults — an override can only lower them | gatewayMs 1,500,000, streamingMs 1,200,000, plainMs 600,000 |
| History | The metrics window and the time-decay tier boundaries and weights | windowMinutes 60 (max 120), tier1Minutes 1, tier2Minutes 5, tier1Weight 10, tier2Weight 3, tier3Weight 1 |
| Sticky | Stable-provider preference: on/off, TTL, hard-switch uptime floor, soft-switch score margin | enabled true, ttlSeconds 3600, uptimeThreshold 85, scoreMargin 0.15 |
| Session | Sticky session routing: on/off, pin TTL, re-pin uptime floor | enabled true, ttlSeconds 3600, uptimeThreshold 85 |
| Provider priorities | Per-provider priority multipliers; set a provider to 0 to disable it for that project | The catalogue's built-in priority for each provider (1 when it has none) |
Per-project routing configuration requires the Enterprise plan. If you'd like to tune routing for your workloads, contact us at contact@llmgateway.io.
Optimized Auto Routing
Auto routing automatically selects the best model for your specific use case without you having to specify a model at all.
Default behaviour
By default, auto routing picks the cheapest of a small built-in set of models that can serve the request: it filters out models whose context window, capabilities (vision, tools, reasoning, structured output) or provider availability do not fit, then selects the cheapest of what remains. Larger prompts skip the smallest model. Requests with audio or document input search every active model that supports that input instead of the built-in set.
# Let LLMGateway choose the optimal model
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Your request here..."}]
}'Free Models Only
When using auto routing, you can restrict the selection to only free models (models with zero input and output pricing) by setting the free_models_only parameter to true:
# Auto route to free models only
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Hello!"}],
"free_models_only": true
}'Adding even a small amount of credits to your account (e.g., $10) will immediately upgrade your free model rate limits from 5 requests per 10 minutes to 20 requests per minute (free-model use still requires a verified email).
The free_models_only parameter only works with auto routing ("model": "auto"). If no free models are available that meet your request requirements,
the API will return an error.
Reasoning models only
Just specify the reasoning_effort value and only a model which supports reasoning will be chosen. This parameter is not specific to the auto model.
# Auto route only to reasoning models
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Hello!"}],
"reasoning_effort": "medium"
}'Exclude Reasoning Models
When using auto routing, you can exclude reasoning models from selection by setting the no_reasoning parameter to true. This is useful when you want faster responses or need to avoid the additional cost and latency of reasoning models:
# Auto route excluding reasoning models
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Hello!"}],
"no_reasoning": true
}'The no_reasoning parameter only works with auto routing ("model": "auto").
If no non-reasoning models are available that meet your request requirements,
the API will return an error.
Auto routing analyzes your payload and automatically chooses between cost-effective models for simple requests and more powerful models for complex or large-context requests.
How It Works
- Request Analysis: The system analyzes your request including message content, context size, and any special parameters
- Model Selection: Based on the analysis, it selects the most appropriate model considering cost, performance, and capabilities
- Transparent Routing: Your request is seamlessly routed to the chosen model and provider
- Optimized Response: You receive the best possible response while maintaining cost efficiency
Auto routing decisions are transparent in your usage logs, so you can always see which model was selected for each request.
Smart Routing
"model": "auto" above is fixed: it always picks from the same built-in set, and its behaviour does not change. Smart routing is a separate model string, "model": "smart", where you choose the candidate models and how they are ranked.
Under Organization settings → Smart Routing you choose up to 30 models from the models catalogue that "model": "smart" may resolve to, plus the classifier that ranks them. Individual projects can override the organization default on their Settings → Routing page; a project without an override inherits it. Owners and organization admins can edit the organization default, project admins the project override.
Smart routing is available to every organization, including pay-as-you-go, while it is in beta. It is not available on DevPass yet. "model": "smart" fails rather than falling back: an organization that cannot use it gets a 403, and one that has not configured it gets a 400 naming the setting — it never degrades quietly into the auto candidate set.
curl -X POST "https://api.llmgateway.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "smart",
"messages": [{"role": "user", "content": "Your request here..."}]
}'The configured list is exhaustive. If no model on it can serve a request — for example an image request against a text-only list — the gateway returns a 400 rather than falling back to a model you did not allow. free_models_only narrows the list to its free models rather than replacing it, so a request parameter cannot route outside what the organization allowed.
Classifiers
- None — pick the cheapest model on the list that can serve the request. This is the default behaviour, restricted to your models.
- Jev (TypeSafe) — rate the request before routing it.
With the Jev classifier, the gateway sorts your models by blended average price and splits them into Low, Medium and High price bands. The dashboard previews that split from catalogue list prices; the gateway ranks only the providers your project can actually use, at your own rates, so the live split can differ. Each request is then classified for task type, output type and difficulty, and served from the matching band; the classifier's own model preference breaks ties inside the band when it is confident enough. The classifier also picks a reasoning effort (low, medium or high), applied when the selected model supports reasoning and clamped to the efforts its provider accepts. An effort you set with reasoning_effort or reasoning.effort always wins. Higher efforts need room: thinking is drawn from the same max_tokens allowance as the answer, so a tight budget steps the effort down (below 16,384 tokens high becomes medium, below 8,192 medium becomes low) rather than risking a response that spends its whole allowance thinking.
The classifier adds one short round trip before the upstream call. It fails open: if it times out or errors, the request is served by the configured fallback model — or the cheapest eligible model on your list when none is set — and the log records the fallback.
What it costs
Smart routing itself carries no platform fee — you pay for the models it selects. A Jev classification is billed at the catalogue rate for jev-1.13.0 on TypeSafe, which is priced on input tokens only, and works out to roughly $0.0001 per call. Each call is recorded as its own log entry against the same organization, project and API key as the request that triggered it, so the amount is visible in your activity feed and usage analytics rather than estimated. The request's own log entry also carries the charge in its smartRouting.classifierCost routing metadata.
The call runs on our credential, so it is billed as credits even for a project using its own provider keys. Nothing is charged for the None classifier, for a verdict reused from a sticky session, or for a classifier call that fails.
Sessions
When a request carries a session id (see sticky session routing) and the project has it enabled, the Jev classifier chooses a model and effort on the session's first turn, and the rest of the session reuses them without another classifier call. Each session keeps its own choice, and the choice expires with the session TTL.
A constant user or prompt_cache_key also counts as a session id, so a client that sends the same value on every request gets the first request's model for all of them. When testing how different prompts are routed, send each one without a session id or with its own x-session-id.
The choice is only reconsidered between turns. A turn starts with a user message after a final answer. Tool results, and user messages sent while the agent is still calling tools, continue the current turn and never change the model. A new turn triggers a recheck in two cases:
- The provider's prompt cache expired. This needs a documented cache lifetime, such as five minutes (or one hour for extended caching) on Claude providers. When a provider's lifetime is unknown, idle time alone never counts as expiry.
- A periodic work scan is due, every fourth turn.
A recheck compares recent instructions and final answers with the work the current choice was made for, and moves only on a confident signal:
- Harder work moves to a stronger model or a higher effort right away.
- Easier or different work moves only when the expected savings for the rest of the session beat the cost of staying by a margin. The estimate counts the check itself, cached versus uncached input prices, the cache rebuild the move causes, and output and reasoning at the new effort. A different model or provider always starts a new cache. An effort-only change is treated as breaking the cache unless the provider documents otherwise. Cost-driven moves are skipped when a price is unknown, and wait a few turns after the previous move.
- Unchanged or unclear work, or a failed check, keeps the current choice. A failed or unsure check is not retried until the next trigger.
If the chosen model stops being available, the stored verdict is re-applied to the remaining candidates rather than triggering a fresh classification. Rechecks are billed like any other classification; reused turns are free.
Responses to "model": "smart" name the resolved model in model. They also carry:
x-llmgateway-smart-model: the resolved model.x-llmgateway-smart-effort: the applied effort.x-llmgateway-smart-change: a short reason, on the turn where the model or effort changed.
Routing metadata
Every smart-routed request made against a configured list records a smartRouting block on its log entry, visible in the request detail view: the classifier used, the eligible and surviving candidate models, the difficulty, task and output type, the classifier's preferred model and confidence, the band that was served, the selected model and effort, the classifier latency and cost, whether the classifier failed open, why it was not consulted (classifierSkipped: single-candidate, compliance or no-credential), whether the configured fallback model served the request, and whether the verdict was reused from earlier in the session. With Jev it also records the probability of each difficulty level (difficultyProbabilities) and of the five most likely candidate models (bestModelProbabilities). Jev returns these typed answers and probabilities rather than a written rationale; the request detail view combines them into a step-by-step Why this model explanation. For a session it also records what triggered the check (initial, reused, mid-turn, cache-expired or scan), the recheck's view of how the work changed, why a recheck kept the current choice, and any switch with its reason and cost estimates.
Routing Savings
The Routing savings card on a project's Analytics page (and org-wide on Organization → Analytics for enterprise admins) shows what auto, smart and dynamic route requests cost compared with routing every one of them to the priciest model the router could have picked.
For each routed request, the gateway prices the request's real token counts (input, cached, output and reasoning) on every candidate model, at the cheapest provider it could have used for that model and your organization's discounts. The most expensive result is the request's baseline; the saving is the baseline minus what the request actually cost, and is never negative. The candidates are:
auto— the built-in model set.smart— the models left on your configured list after filtering for the request.- Dynamic routes — every catalogue model node in the route version that served the request. Custom-provider targets have no catalogue price and are left out.
Cache hits, retried attempts and requests without token usage are excluded. Each request detail view shows its baseline under Cost Information.
The saving is an estimate: a different model would have produced a different number of output tokens. Requests from the 30 days before this feature launched were priced once, retroactively, at the catalogue prices in effect at launch.
Best Practices
For Development
- Use specific model names during development and testing
- Leverage auto routing for production workloads to optimize costs
For Production
- Use auto routing (
"model": "auto") for the best balance of cost and performance, or smart routing ("model": "smart") to route across models you choose - Monitor your usage patterns through the dashboard to understand routing decisions
- Set up provider keys for multiple providers to maximize routing options
For Cost Optimization
- Let auto routing handle model selection to automatically use the most cost-effective options, or smart routing to spend more only on the requests that need it
- Use model IDs without provider prefixes to always get the cheapest available provider
- Monitor your usage analytics to track cost savings from intelligent routing
How is this guide?
Last updated on