> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mithunai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Routing and failover in the AI Gateway

> How the Gateway chooses a model and a credential for each request, what happens when an upstream fails, and the routing policy a platform operator can set.

The Gateway chooses one `(provider, model, credential)` for every request and retries on a different one when an upstream fails. This page describes how that choice is made and how it is configured.

<Note>
  Routing policy is **platform** administration, not customer administration. The Gateway dashboard
  is the Platform Super Admin's surface, and the settings below are deployment-wide ([Organisations
  and roles](/concepts/organizations-and-roles)). A client steers a single request through the
  `model` field instead — described first, because that part needs no administration at all.
</Note>

## Steering one request: the `model` field

Send `model: "auto"` and the request follows the deployment's active fallback chain. A suffix steers that one request without changing anything:

| `model` | What it ranks |
| - | - |
| `auto` | the active fallback chain, in its configured order |
| `auto:smart` | highest-capability models first |
| `auto:fast` | measured speed first — throughput blended with time-to-first-byte |
| `auto:reliable` | recent success rate first |
| `auto:balanced` | the default blend: reliability first, speed and capability sharing the rest |
| `auto:cheap` | budget-leaning; currently the same blend as `balanced` |
| `auto:<chain-name>` | a named chain, instead of the active one |

The axis suffixes rank **every enabled model** and ignore chain order. Common synonyms resolve too (`auto:fastest`, `auto:speed`, `auto:smartest`, `auto:cheapest`, `auto:budget`), and the whole string is case-insensitive.

```bash theme={"theme":{"light":"github-light","dark":"github-dark"}}
curl "$GATEWAY_URL/v1/chat/completions" \
  --header "Authorization: Bearer $GATEWAY_API_KEY" \
  --header "Content-Type: application/json" \
  --data '{
    "model": "auto:fast",
    "messages": [{"role": "user", "content": "Summarise this release note in one line."}]
  }'
```

The response carries an `x-routed-via` header naming what actually served the request.

## Named chains

A named chain is a hand-picked ordered list of models. Every chain is advertised to clients in `GET /v1/models` as the model id `auto:<name>`, so two tools can use two different chains through the same key with no dashboard change and no key rotation. An unknown chain name is a `400`, never a silent fall back to the active chain.

A new chain starts **empty** and opts out of catalogue backfill, which is the point of naming one: the operator picks these models, in this order, and the Gateway stops guessing. A chain can opt back in, in which case each newly discovered model is appended at the next priority, inheriting that model's own enabled state. An empty chain refuses with a `400` rather than quietly routing over the whole catalogue.

### Auto-ordering a chain

Rather than dragging rows, a chain can be reordered by one of three presets: **Smartest** (capability tier, then rank within the tier), **Fastest** (measured speed rank) or **Biggest budget** (monthly free-tier allowance, largest first). Each writes priority only — a preset never switches a model on or off, and a model the chain has not named yet is written in switched **off**, the state it was displayed in.

## How a model is scored

Ranking is not round-robin. Three normalised axes are combined with weights that sum to one, and two guardrail multipliers then scale the result:

```
base      = w_reliability·reliability + w_speed·speed + w_capability·capability
effective = base × headroom factor × rate-limit factor
```

The weights come from a selectable strategy — `balanced` (the default, 0.5 / 0.25 / 0.25), `smartest`, `fastest`, `reliable`, `custom` or `priority`. Reliability is learned from live traffic rather than configured. The two multipliers are always on: a model close to its free-tier allowance, or carrying a live rate-limit penalty, is held back below `1.0`. Both multipliers have operator-settable thresholds — see [Guardrails and settings](/gateway/guardrails).

Choosing *which credential* to use within a model is a separate decision (`auto` or least-remaining-quota), deliberately independent of model ranking so the two do not interfere.

## Conditional routing

Beyond one active chain, an operator can store an ordered list of rules, each a condition plus the chain or axis to route to when it matches. Rules apply **only** when the request did not name a target: an explicit `auto:<name>` always wins, so a rule can never override a deliberate client choice.

A condition reads two closed namespaces and nothing else. `params.*` is derived by the Gateway from the request it is already serving:

| Key | Value |
| - | - |
| `params.endpoint` | which surface the request arrived on |
| `params.stream` | whether the client asked for a stream |
| `params.tools` | whether the request declared tools |
| `params.task` | the derived task type, `code` or `chat` |
| `params.agent` | the detected client, or `unknown` |
| `params.model` | the model string the client sent |

`metadata.*` reads labels the client declares in a request header (the header name is the `ROUTING_METADATA_HEADER` constant in `server/src/services/routing-rules.ts`). Client-declared labels are treated as untrusted: the worst a forged value can do is select a different chain **within the caller's own organisation**, which the caller could already do by sending `model: "auto:<name>"`. No condition key can name an organisation, a credential or a key — the namespaces are closed.

Conditions use a JSON query form with the comparison operators `$eq`, `$ne`, `$gt`, `$gte`, `$lt`, `$lte`, `$in`, `$nin` and `$regex`, plus `$and` and `$or`:

```json theme={"theme":{"light":"github-light","dark":"github-dark"}}
{
  "enabled": true,
  "rules": [
    { "when": { "params.task": "code" }, "then": "smart" },
    { "when": { "params.stream": true, "params.tools": false }, "then": "fast" }
  ],
  "default": null
}
```

A rule set is validated when it is stored, and an unknown operator or an unreachable key is refused there with a message. Evaluation on the request path never throws: a rule that somehow went stale fails to match rather than turning a live request into a `500`, and a corrupt stored rule set reads as disabled rather than stopping routing. Nesting depth, clause count, rule count and regular-expression length are all bounded, so a stored rule cannot cost an unbounded walk per request.

A rule naming a chain the calling organisation does not own resolves to nothing, and the request falls through to the chain it would have used anyway.

Read and write the rule set with `GET` and `PUT /api/fallback/routing-rules`.

## What happens when an upstream fails

Failover is one shared loop for every chat-shaped surface, so the OpenAI-, Responses- and Anthropic-shaped endpoints behave identically.

* **Hops are capped.** A request makes at most 20 failover attempts.
* **There is a wall-clock budget.** By default the loop stops *starting* further retries after 45 seconds and cancels an attempt still waiting on its first byte. The first attempt always runs, and so does the first retry, so one failover hop is always structurally possible even if the first attempt consumed the whole budget. Set it to `0` to disable it.
* **A failing credential is benched.** A retryable upstream failure — `401`, `429`, `5xx`, an empty stream, a timeout — puts that credential on a cooldown, and a `Retry-After` from the provider is honoured.
* **A failing model is benched across every credential.** Three retryable failures inside a 15-minute sliding window bench the whole model for 10 minutes on *every* credential that could serve it, because the window counts across credentials; benching only the one that failed last leaves every sibling serving the same sick model. Recovery is automatic — a background probe re-validates a benched credential partway through, and a model still sick simply re-trips the streak.
* **A client cancelling is not a failure.** A client-side abort returns without recording anything against the model or the credential.
* **A per-model breaker is available.** An operator can additionally stop trying a model after N consecutive upstream failures; `0` disables it.

### Degraded mode

When a large share of providers fails at once, probing and exploration only burn retry budget on dead routes. A state machine tracks the healthy-provider ratio from the scheduled health pass and flips the deployment into a degraded state once the ratio stays below its threshold for long enough. Entry and exit both need the ratio to hold, and exit needs it for longer, so one bad pass cannot flap the state.

| Setting | Default | Meaning |
| - | - | - |
| `DEGRADED_HEALTHY_RATIO` | `0.5` | share of providers that must have at least one usable credential |
| `DEGRADED_MIN_PROVIDERS` | `3` | below this many providers the state is not evaluated, so a single-provider deployment never flaps |
| `DEGRADED_ENTRY_GRACE_MS` | `60000` | how long the ratio must stay low before entering |
| `DEGRADED_EXIT_GRACE_MS` | `120000` | how long it must stay healthy before leaving |

While degraded, exploration is switched off and the router sticks to the scored order of the providers that remain. An unprobed credential counts as healthy — it is treated as usable until a probe says otherwise. The health endpoint reports the state and when it was entered, so an operator can see it.

## Where this is configured

| Setting | Endpoint |
| - | - |
| Routing strategy and weights | `GET` / `PUT /api/fallback/routing` |
| Conditional routing rules | `GET` / `PUT /api/fallback/routing-rules` |
| The chain itself | `GET` / `PUT /api/fallback` |
| Auto-order a chain | `POST /api/fallback/sort/:preset` — `intelligence`, `speed` or `budget` |
| Guardrail thresholds | see [Guardrails and settings](/gateway/guardrails) |

All of these sit behind the Gateway dashboard session. They are not part of the [MITHUNAI HTTP API](/channels/api), which is the customer-facing surface.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.