> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mithunai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# MITHUNAI AI Gateway

> Universal model routing, 200+ foundation models, automatic fallback chains, and unified token budget governance across your enterprise.

MITHUNAI AI Gateway (`gateway.mithunai.com`) is an enterprise-grade model routing and governance proxy that unifies access to over 200+ LLMs, vision, embedding, audio, and video models through a single OpenAI-compatible interface.

Instead of hardcoding vendor SDKs or scattering provider API keys across team repositories, all downstream applications, agents, and conversational interfaces call the gateway. The gateway manages credentials, enforces budgets, optimizes latency and cost, and executes zero-downtime fallback chains when upstream providers degrade.

<Columns cols={3}>
  <Card title="200+ Foundation Models" icon="cpu">
    Connect to OpenAI, Anthropic Claude, Google Gemini, DeepSeek, Mistral, Groq, Bedrock, and local
    Ollama instances.
  </Card>

  <Card title="Intelligent Routing" icon="route">
    Route requests dynamically based on real-time latency, provider cost, model capability tiers, or
    custom weighted scores.
  </Card>

  <Card title="Zero-Downtime Fallbacks" icon="life-buoy">
    Define ordered fallback chains. When an upstream provider rate limits (429) or errors (5xx),
    requests fail over instantly.
  </Card>
</Columns>

***

## Universal OpenAI-Compatible API

MITHUNAI Gateway exposes standard OpenAI-compatible endpoints, allowing drop-in compatibility with existing codebases, LangChain, LlamaIndex, LiteLLM, and developer tools:

```bash theme={"theme":{"light":"github-light","dark":"github-dark"}}
curl -X POST https://gateway.mithunai.com/v1/chat/completions \
  -H "Authorization: Bearer mth_gw_live_xxxxxxxxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-3-7-sonnet",
    "messages": [
      {"role": "system", "content": "You are an enterprise AI assistant."},
      {"role": "user", "content": "Analyze quarterly platform uptime."}
    ],
    "stream": true
  }'
```

Supported modality routes:

* **Chat & Reasoning**: `/v1/chat/completions` (streaming SSE supported)
* **Embeddings**: `/v1/embeddings` (vector representations for semantic indexing)
* **Image Generation**: `/v1/images/generations`
* **Audio & Speech**: `/v1/audio/transcriptions` and `/v1/audio/speech`
* **Model Discovery**: `/v1/models`

***

## Multi-Modality Hub

The Gateway dashboard provides dedicated workspaces for exploring and benchmarking models across all modalities:

1. **Interactive Playground**: Test prompts across chat models, adjust temperature, top\_p, and inspect raw streaming tokens in real time.
2. **Fusion & Ensemble Routing**: Send prompts concurrently to multiple models to compare reasoning depth and output diversity.
3. **Embeddings Inspector**: Evaluate vector dimensions and similarity metrics across proprietary and open-source embedding models.
4. **Media Generation**: Test image, audio, and video generation endpoints with unified credentials.

***

## Intelligent Routing & Fallback Strategies

Configure fallback rules per model alias so mission-critical services never experience downtime:

```mermaid theme={"theme":{"light":"github-light","dark":"github-dark"}}
flowchart LR
    A["Application Request"] --> B["MITHUNAI Gateway"]
    B --> C{"Primary Route"}
    C -- "Healthy" --> D["Claude 3.7 Sonnet"]
    C -- "429 / 5xx / Timeout" --> E{"Fallback 1"}
    E -- "Healthy" --> F["GPT-4o"]
    E -- "Failover" --> G["Gemini 2.5 Pro"]
```

Routing algorithms supported:

* **Priority Cascade**: Attempt the primary model first, seamlessly cascading down the fallback sequence if an error occurs.
* **Latency Optimization**: Direct traffic to the provider endpoint currently demonstrating the lowest time-to-first-token (TTFT).
* **Cost Minimization**: Route non-critical workloads (e.g. summarization, classification) to the most cost-effective tier meeting the requirement.
* **Load Balancing**: Distribute high-volume traffic across multiple upstream provider accounts or regions.

***

## Virtual API Keys & Budget Governance

Never expose master provider API keys to developer teams or production microservices:

* **Virtual Keys**: Generate scoped keys (`mth_gw_...`) with fine-grained access policies per department or service.
* **Hard & Soft Spend Limits**: Set monthly dollar caps or token quotas. Notify administrators when thresholds (e.g., 80%) are reached, and reject calls when limits are breached.
* **Model Scoping**: Restrict specific keys to sanctioned model lists (e.g., only allow internal models or specific compliant providers).
* **Detailed Audit Logs**: Capture token usage, latency metrics, prompt tokens, completion tokens, and exact cost attribution per request.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.