(provider, model, credential) for every request and retries on a different one when an upstream fails. This page describes how that choice is made and how it is configured.
Routing policy is platform administration, not customer administration. The Gateway dashboard
is the Platform Super Admin’s surface, and the settings below are deployment-wide (Organisations
and roles). A client steers a single request through the
model field instead — described first, because that part needs no administration at all.Steering one request: the model field
Send model: "auto" and the request follows the deployment’s active fallback chain. A suffix steers that one request without changing anything:
The axis suffixes rank every enabled model and ignore chain order. Common synonyms resolve too (
auto:fastest, auto:speed, auto:smartest, auto:cheapest, auto:budget), and the whole string is case-insensitive.
x-routed-via header naming what actually served the request.
Named chains
A named chain is a hand-picked ordered list of models. Every chain is advertised to clients inGET /v1/models as the model id auto:<name>, so two tools can use two different chains through the same key with no dashboard change and no key rotation. An unknown chain name is a 400, never a silent fall back to the active chain.
A new chain starts empty and opts out of catalogue backfill, which is the point of naming one: the operator picks these models, in this order, and the Gateway stops guessing. A chain can opt back in, in which case each newly discovered model is appended at the next priority, inheriting that model’s own enabled state. An empty chain refuses with a 400 rather than quietly routing over the whole catalogue.
Auto-ordering a chain
Rather than dragging rows, a chain can be reordered by one of three presets: Smartest (capability tier, then rank within the tier), Fastest (measured speed rank) or Biggest budget (monthly free-tier allowance, largest first). Each writes priority only — a preset never switches a model on or off, and a model the chain has not named yet is written in switched off, the state it was displayed in.How a model is scored
Ranking is not round-robin. Three normalised axes are combined with weights that sum to one, and two guardrail multipliers then scale the result:balanced (the default, 0.5 / 0.25 / 0.25), smartest, fastest, reliable, custom or priority. Reliability is learned from live traffic rather than configured. The two multipliers are always on: a model close to its free-tier allowance, or carrying a live rate-limit penalty, is held back below 1.0. Both multipliers have operator-settable thresholds — see Guardrails and settings.
Choosing which credential to use within a model is a separate decision (auto or least-remaining-quota), deliberately independent of model ranking so the two do not interfere.
Conditional routing
Beyond one active chain, an operator can store an ordered list of rules, each a condition plus the chain or axis to route to when it matches. Rules apply only when the request did not name a target: an explicitauto:<name> always wins, so a rule can never override a deliberate client choice.
A condition reads two closed namespaces and nothing else. params.* is derived by the Gateway from the request it is already serving:
metadata.* reads labels the client declares in a request header (the header name is the ROUTING_METADATA_HEADER constant in server/src/services/routing-rules.ts). Client-declared labels are treated as untrusted: the worst a forged value can do is select a different chain within the caller’s own organisation, which the caller could already do by sending model: "auto:<name>". No condition key can name an organisation, a credential or a key — the namespaces are closed.
Conditions use a JSON query form with the comparison operators $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin and $regex, plus $and and $or:
500, and a corrupt stored rule set reads as disabled rather than stopping routing. Nesting depth, clause count, rule count and regular-expression length are all bounded, so a stored rule cannot cost an unbounded walk per request.
A rule naming a chain the calling organisation does not own resolves to nothing, and the request falls through to the chain it would have used anyway.
Read and write the rule set with GET and PUT /api/fallback/routing-rules.
What happens when an upstream fails
Failover is one shared loop for every chat-shaped surface, so the OpenAI-, Responses- and Anthropic-shaped endpoints behave identically.- Hops are capped. A request makes at most 20 failover attempts.
- There is a wall-clock budget. By default the loop stops starting further retries after 45 seconds and cancels an attempt still waiting on its first byte. The first attempt always runs, and so does the first retry, so one failover hop is always structurally possible even if the first attempt consumed the whole budget. Set it to
0to disable it. - A failing credential is benched. A retryable upstream failure —
401,429,5xx, an empty stream, a timeout — puts that credential on a cooldown, and aRetry-Afterfrom the provider is honoured. - A failing model is benched across every credential. Three retryable failures inside a 15-minute sliding window bench the whole model for 10 minutes on every credential that could serve it, because the window counts across credentials; benching only the one that failed last leaves every sibling serving the same sick model. Recovery is automatic — a background probe re-validates a benched credential partway through, and a model still sick simply re-trips the streak.
- A client cancelling is not a failure. A client-side abort returns without recording anything against the model or the credential.
- A per-model breaker is available. An operator can additionally stop trying a model after N consecutive upstream failures;
0disables it.
Degraded mode
When a large share of providers fails at once, probing and exploration only burn retry budget on dead routes. A state machine tracks the healthy-provider ratio from the scheduled health pass and flips the deployment into a degraded state once the ratio stays below its threshold for long enough. Entry and exit both need the ratio to hold, and exit needs it for longer, so one bad pass cannot flap the state.
While degraded, exploration is switched off and the router sticks to the scored order of the providers that remain. An unprobed credential counts as healthy — it is treated as usable until a probe says otherwise. The health endpoint reports the state and when it was entered, so an operator can see it.
Where this is configured
All of these sit behind the Gateway dashboard session. They are not part of the MITHUNAI HTTP API, which is the customer-facing surface.