> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mithunai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Gateway guardrails and settings

> The deployment-wide knobs in Settings → Gateway: output cap, per-request token budget, failover breaker, headroom thresholds, task-type weight shift and the response cache.

Seven settings bound what a single request may cost and how hard the router works before giving up. They are deployment-wide, they take effect on the **next request** — none of them needs a restart — and they are reachable both from **Settings → Gateway** in the dashboard and over the dashboard API.

<Note>
  These are **platform** administration (Platform Super Admin), not customer administration, and
  they are not part of the customer-facing [MITHUNAI HTTP API](/channels/api).
</Note>

## Bounding one request

### Output token cap

Clamps every request's `max_tokens` before it reaches a provider. Three states: **Off** passes the caller's value straight through, **Auto** clamps to 32,768, and a custom integer clamps to that value.

This is the cheapest protection against a client that forgets to set a limit at all, because the cost of an unbounded completion is paid at the provider, not here.

`GET` / `PUT /api/settings/output-limit` — `{"mode": "off" | "auto" | <integer>}`.

### Per-request token budget

Refuses a request outright when its estimated prompt plus output exceeds a number of tokens. `0` disables the check.

The cap above *shortens* an answer; this one *declines* the question. Use it when a single oversized request is worse than a refused one.

`GET` / `PUT /api/settings/guardrails` — `requestMaxTokensBudget`.

## Bounding failover

### Failover circuit breaker

How many consecutive upstream failures a model may produce before the router stops trying it and falls through to the next. `0` disables the breaker.

This is in addition to, not instead of, the automatic benching described in [Routing and failover](/gateway/routing-and-fallbacks) — that benches a model after three failures in a sliding window regardless of this setting. The breaker is the knob for a deployment that wants to give up sooner.

`GET` / `PUT /api/settings/guardrails` — `maxConsecutiveUpstreamFails`.

## Shaping the router

Two of the three scoring axes are learned or measured. These two settings change how the guardrail multipliers and the task bias behave. Both accept a decimal between 0 and 1, and both accept `null` to clear back to the built-in default.

### Headroom ramp start

The remaining-allowance share at which the router begins **proactively demoting** a model — `0.2` means "start easing off at 20% remaining". Ramping down before an allowance runs out is what turns a hard `429` cliff into a gradual shift of traffic.

### Headroom score floor

The score a model keeps once its allowance is exhausted, as a share of its normal score. A floor above zero keeps an exhausted model as a last resort rather than removing it entirely, which matters on a deployment where everything is exhausted at once.

`GET` / `PUT /api/settings/headroom` — `{"rampStart": …, "floor": …}`.

### Task-type weight shift

When a request's task type can be derived — code or chat — this is the share of one routing axis moved onto the other: **code leans on capability, chat leans on speed**. `0` disables the bias; clearing it uses the built-in default.

`GET` / `PUT /api/settings/task-weight-share` — `{"share": 0..1 | null}`.

## The response cache

Off by default. When it is on, an identical request is served from memory instead of spending a provider call.

What "identical" means is strict, and deliberately so:

* **Exact match only.** The key is a SHA-256 over the canonicalised request. There is no embedding or fuzzy matching, so a near-miss can never return a different prompt's answer — one differing token is a miss.
* **The key is the request, not the route.** Any model's good answer to an identical request is a valid hit, which is what makes the cache worth having for auto-routed traffic.
* **A hit cannot cross organisations.** The organisation is hashed into the cache key, so two organisations asking the identical question never compute the same key. That is the isolation; the organisation recorded on the row is for attribution and for purging, not for the control.
* **Temperature-gated.** A high-temperature request is asking for variety, so replaying one frozen answer would defeat it. Only requests with no temperature, or a low one, are cached.
* **Bounded.** Entries live in a size-capped LRU with a TTL, so the cache cannot grow without bound.

<Warning>
  **Persistence is a privacy decision, not a performance one.** With the cache on, entries are also
  written through to disk so that a restart does not throw away the day's savings — which means
  **plaintext model responses land on disk**. Turn persistence off for a memory-only cache if that
  is not acceptable for your deployment, and read [Security](/security) for what the platform does
  with content generally.
</Warning>

`GET /api/cache/stats` reports entries, TTL and hit rate; `PUT /api/cache/config` turns it on or off; `DELETE /api/cache` flushes it. A per-request header can override the deployment setting either way.

## All seven, at a glance

| Setting | Default | Endpoint |
| - | - | - |
| Output token cap | off | `/api/settings/output-limit` |
| Per-request token budget | `0` (disabled) | `/api/settings/guardrails` |
| Failover circuit breaker | `0` (disabled) | `/api/settings/guardrails` |
| Headroom ramp start | built-in default | `/api/settings/headroom` |
| Headroom score floor | built-in default | `/api/settings/headroom` |
| Task-type weight shift | built-in default | `/api/settings/task-weight-share` |
| Response cache | off | `/api/cache/config` |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.