Skip to main content
Seven settings bound what a single request may cost and how hard the router works before giving up. They are deployment-wide, they take effect on the next request — none of them needs a restart — and they are reachable both from Settings → Gateway in the dashboard and over the dashboard API.
These are platform administration (Platform Super Admin), not customer administration, and they are not part of the customer-facing MITHUNAI HTTP API.

Bounding one request

Output token cap

Clamps every request’s max_tokens before it reaches a provider. Three states: Off passes the caller’s value straight through, Auto clamps to 32,768, and a custom integer clamps to that value. This is the cheapest protection against a client that forgets to set a limit at all, because the cost of an unbounded completion is paid at the provider, not here. GET / PUT /api/settings/output-limit — {"mode": "off" | "auto" | <integer>}.

Per-request token budget

Refuses a request outright when its estimated prompt plus output exceeds a number of tokens. 0 disables the check. The cap above shortens an answer; this one declines the question. Use it when a single oversized request is worse than a refused one. GET / PUT /api/settings/guardrails — requestMaxTokensBudget.

Bounding failover

Failover circuit breaker

How many consecutive upstream failures a model may produce before the router stops trying it and falls through to the next. 0 disables the breaker. This is in addition to, not instead of, the automatic benching described in Routing and failover — that benches a model after three failures in a sliding window regardless of this setting. The breaker is the knob for a deployment that wants to give up sooner. GET / PUT /api/settings/guardrails — maxConsecutiveUpstreamFails.

Shaping the router

Two of the three scoring axes are learned or measured. These two settings change how the guardrail multipliers and the task bias behave. Both accept a decimal between 0 and 1, and both accept null to clear back to the built-in default.

Headroom ramp start

The remaining-allowance share at which the router begins proactively demoting a model — 0.2 means “start easing off at 20% remaining”. Ramping down before an allowance runs out is what turns a hard 429 cliff into a gradual shift of traffic.

Headroom score floor

The score a model keeps once its allowance is exhausted, as a share of its normal score. A floor above zero keeps an exhausted model as a last resort rather than removing it entirely, which matters on a deployment where everything is exhausted at once. GET / PUT /api/settings/headroom — {"rampStart": …, "floor": …}.

Task-type weight shift

When a request’s task type can be derived — code or chat — this is the share of one routing axis moved onto the other: code leans on capability, chat leans on speed. 0 disables the bias; clearing it uses the built-in default. GET / PUT /api/settings/task-weight-share — {"share": 0..1 | null}.

The response cache

Off by default. When it is on, an identical request is served from memory instead of spending a provider call. What “identical” means is strict, and deliberately so:
  • Exact match only. The key is a SHA-256 over the canonicalised request. There is no embedding or fuzzy matching, so a near-miss can never return a different prompt’s answer — one differing token is a miss.
  • The key is the request, not the route. Any model’s good answer to an identical request is a valid hit, which is what makes the cache worth having for auto-routed traffic.
  • A hit cannot cross organisations. The organisation is hashed into the cache key, so two organisations asking the identical question never compute the same key. That is the isolation; the organisation recorded on the row is for attribution and for purging, not for the control.
  • Temperature-gated. A high-temperature request is asking for variety, so replaying one frozen answer would defeat it. Only requests with no temperature, or a low one, are cached.
  • Bounded. Entries live in a size-capped LRU with a TTL, so the cache cannot grow without bound.
Persistence is a privacy decision, not a performance one. With the cache on, entries are also written through to disk so that a restart does not throw away the day’s savings — which means plaintext model responses land on disk. Turn persistence off for a memory-only cache if that is not acceptable for your deployment, and read Security for what the platform does with content generally.
GET /api/cache/stats reports entries, TTL and hit rate; PUT /api/cache/config turns it on or off; DELETE /api/cache flushes it. A per-request header can override the deployment setting either way.

All seven, at a glance

Last modified on October 5, 2026