Provider credentials and the catalogue are platform administration, managed from the Gateway
dashboard by the Platform Super Admin. A customer-facing assistant chooses from the models the
deployment offers — see Models and providers.
What a client sees
GET /v1/models returns the catalogue in the OpenAI shape, alongside the Claude-family discovery ids and one entry per named chain as auto:<name>. That is the authoritative list for a deployment; this page deliberately does not reproduce one, because the answer is per deployment and changes whenever a credential does.
What the catalogue records per model
Routing needs comparable facts about models from unrelated providers, so each row carries:- a capability tier — Frontier, Large, Medium or Small — which is the only cross-provider comparison that holds, plus a rank within the provider’s own catalogue used as a tiebreaker inside a tier;
- a measured speed rank, written back from observed throughput and time-to-first-byte rather than configured;
- the monthly allowance published for the model on its free tier, stored as the label the provider states (
~120M,~3M (1k credits),credits-based), and a daily cap where the provider gives one; - capability flags for tool calling and vision, where the provider or the operator declares them.
Free-tier allowances are pooled, not summed
Many models on one provider share a single free allowance, so adding up every model’s published label double-counts badly. The Gateway groups models into quota pools and counts one documented budget per pool — the largest label seen in it. Because a published allowance is per account, a pool is scaled by the number of usable credentials you hold for that provider: two usable accounts really are twice the pool.GET /api/free-tier returns those pools, each marked as:
Where a provider reports live quota in its responses, the pool also carries remaining, limit and the next reset, summed across the credentials in it and labelled with the metric the provider actually used — tokens, credits, neurons or requests. Tokens are preferred when several are available, because a rate-limit counter read as a token budget is the wrong number under the right heading. A model that is switched off in the chain still counts toward its pool, because it draws on the same provider allowance; it is marked, not dropped.
This is the same data the dashboard’s free-tier panel and its per-model usage bar read, so the two never disagree about the same pool.
What the cost figures mean
The Gateway does not set prices, bill anyone, or mark anything up. You pay your providers under your own agreements with them; the Gateway holds your credentials and routes to them. What it does hold is a paid-equivalent reference table: for each model, what the same model — or its nearest equivalent — costs per million input and output tokens on a paid API. Its only use is the estimated savings figure in analytics, so that a free-tier call is credited at a realistic price rather than priced as though every token were a frontier model’s. Two things follow, and both matter if you report that number onward:- It is a snapshot, not a live feed. The table was taken from public pricing on 2026-06-05, with official API prices used for closed models. It does not track price changes, and nothing in the Gateway re-fetches it.
- Some models have no paid equivalent at all — preview and stealth models — and analytics falls back to a modest default for those.
Self-hosted and private models
A model does not have to be a hosted API.- Ollama is served through its own compatibility surface, so a local runtime can be reached through the same Gateway.
- Any OpenAI-compatible endpoint can be registered as a custom provider. An endpoint is identified by its base URL and holds its own pool of credentials, so two endpoints offering the same model id each keep their own row — their own enabled flag, their own health and their own measured speed. One of them being broken does not disable the other, which was the whole reason for keying identity on the endpoint.
- Models for a custom endpoint can be declared alongside its credential, including tool and vision support.