These are platform administration (Platform Super Admin), not customer administration, and
they are not part of the customer-facing MITHUNAI HTTP API.
Bounding one request
Output token cap
Clamps every request’smax_tokens before it reaches a provider. Three states: Off passes the caller’s value straight through, Auto clamps to 32,768, and a custom integer clamps to that value.
This is the cheapest protection against a client that forgets to set a limit at all, because the cost of an unbounded completion is paid at the provider, not here.
GET / PUT /api/settings/output-limit — {"mode": "off" | "auto" | <integer>}.
Per-request token budget
Refuses a request outright when its estimated prompt plus output exceeds a number of tokens.0 disables the check.
The cap above shortens an answer; this one declines the question. Use it when a single oversized request is worse than a refused one.
GET / PUT /api/settings/guardrails — requestMaxTokensBudget.
Bounding failover
Failover circuit breaker
How many consecutive upstream failures a model may produce before the router stops trying it and falls through to the next.0 disables the breaker.
This is in addition to, not instead of, the automatic benching described in Routing and failover — that benches a model after three failures in a sliding window regardless of this setting. The breaker is the knob for a deployment that wants to give up sooner.
GET / PUT /api/settings/guardrails — maxConsecutiveUpstreamFails.
Shaping the router
Two of the three scoring axes are learned or measured. These two settings change how the guardrail multipliers and the task bias behave. Both accept a decimal between 0 and 1, and both acceptnull to clear back to the built-in default.
Headroom ramp start
The remaining-allowance share at which the router begins proactively demoting a model —0.2 means “start easing off at 20% remaining”. Ramping down before an allowance runs out is what turns a hard 429 cliff into a gradual shift of traffic.
Headroom score floor
The score a model keeps once its allowance is exhausted, as a share of its normal score. A floor above zero keeps an exhausted model as a last resort rather than removing it entirely, which matters on a deployment where everything is exhausted at once.GET / PUT /api/settings/headroom — {"rampStart": …, "floor": …}.
Task-type weight shift
When a request’s task type can be derived — code or chat — this is the share of one routing axis moved onto the other: code leans on capability, chat leans on speed.0 disables the bias; clearing it uses the built-in default.
GET / PUT /api/settings/task-weight-share — {"share": 0..1 | null}.
The response cache
Off by default. When it is on, an identical request is served from memory instead of spending a provider call. What “identical” means is strict, and deliberately so:- Exact match only. The key is a SHA-256 over the canonicalised request. There is no embedding or fuzzy matching, so a near-miss can never return a different prompt’s answer — one differing token is a miss.
- The key is the request, not the route. Any model’s good answer to an identical request is a valid hit, which is what makes the cache worth having for auto-routed traffic.
- A hit cannot cross organisations. The organisation is hashed into the cache key, so two organisations asking the identical question never compute the same key. That is the isolation; the organisation recorded on the row is for attribution and for purging, not for the control.
- Temperature-gated. A high-temperature request is asking for variety, so replaying one frozen answer would defeat it. Only requests with no temperature, or a low one, are cached.
- Bounded. Entries live in a size-capped LRU with a TTL, so the cache cannot grow without bound.
GET /api/cache/stats reports entries, TTL and hit rate; PUT /api/cache/config turns it on or off; DELETE /api/cache flushes it. A per-request header can override the deployment setting either way.