# OmniRoute

> Never stop coding. Free AI gateway: one endpoint, 160+ providers (50+ free), connect Claude Code, Codex, Cursor, Cline & Copilot to FREE Claude/GPT/Gemini. RTK+Caveman stacked compression saves 15-95% tokens, smart auto-fallback, MCP/A2A, multimodal APIs, Desktop/PWA.

- **Type:** MCP server
- **Install:** `agentstack add mcp-diegosouzapw-omniroute`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [diegosouzapw](https://agentstack.voostack.com/s/diegosouzapw)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [diegosouzapw](https://github.com/diegosouzapw)
- **Source:** https://github.com/diegosouzapw/OmniRoute
- **Website:** https://omniroute.online

## Install

```sh
agentstack add mcp-diegosouzapw-omniroute
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# 🚀 OmniRoute — The Free AI Gateway

### Never stop coding. Connect every AI tool to **237 providers** — **50+ free** — through one endpoint.

**Plug Claude Code, Codex, Cursor, Cline, Copilot & Antigravity into FREE Claude / GPT / Gemini. Auto-fallback.**

**RTK + Caveman compression saves 15–95% tokens. Never hit limits.**

**~1.6B documented free tokens/month** — up to **~2.1B in your first month** with signup credits — aggregated across the free tiers, plus a long tail of permanently-free, no-cap providers, and the compression above stretches every one further. ([how we count →](docs/reference/FREE_TIERS.md#tldr--how-much-free-inference-does-omniroute-actually-aggregate))

[](#-231-ai-providers--50-free)
[](#-231-ai-providers--50-free)
[](docs/reference/FREE_TIERS.md)
[](#%EF%B8%8F-save-1595-tokens--automatically)
[](#-combos--the-flagship)
[](#-quick-start)

### 💬 Join the community

[](https://discord.gg/EkzRkpzKYt)
[](https://t.me/omnirouteOficial)
[](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t)
[](https://chat.whatsapp.com/BTGJXIyjeNIIgExvTMGGhI)

**Questions, provider tips, roadmap & support → [Discord](https://discord.gg/EkzRkpzKYt) · [Telegram](https://t.me/omnirouteOficial) · WhatsApp [🌍 Global](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) / [🇧🇷 Brasil](https://chat.whatsapp.com/BTGJXIyjeNIIgExvTMGGhI)**

[](https://www.npmjs.com/package/omniroute)
[](LICENSE)
[](package.json)
[](https://github.com/diegosouzapw/OmniRoute)

[](https://www.npmjs.com/package/omniroute)

[](https://hub.docker.com/r/diegosouzapw/omniroute)

[](https://omniroute.online)

[**🚀 Quick Start**](#-quick-start) • [**🎯 Combos**](#-combos--the-flagship) • [**🌐 Providers**](#-231-ai-providers--50-free) • [**🔌 CLI & MCP**](#-full-cli--a2a--mcp) • [**🗜️ Compression**](#%EF%B8%8F-save-1595-tokens--automatically) • [**🌍 Website**](https://omniroute.online)

[💥 The Promise](#-the-promise) • [🤔 Why](#-why-omniroute) • [🏆 What Sets Apart](#-what-sets-omniroute-apart) • [🤖 Compatible CLIs](#-compatible-clis--coding-agents) • [🖥️ Where It Runs](#%EF%B8%8F-where-omniroute-runs--anywhere) • [🔒 Private](#-private--local-first) • [🎬 In Action](#-omniroute-in-action) • [📚 Explore More](#-explore-more) • [📧 Support](#-support--community)

 🌐 Available in 41+ languages
 
  
    🇺🇸
    🇧🇷
    🇪🇸
    🇫🇷
    🇮🇹
    🇷🇺
    🇨🇳
    🇹🇼
    🇩🇪
    🇯🇵
    🇰🇷
    🇮🇳
  
  
    🇹🇭
    🇻🇳
    🇮🇩
    🇲🇾
    🇵🇭
    🇸🇦
    🇮🇱
    🇦🇿
    🇺🇦
    🇵🇱
    🇨🇿
  
  
    🇳🇱
    🇧🇬
    🇩🇰
    🇫🇮
    🇳🇴
    🇸🇪
    🇭🇺
    🇷🇴
    🇸🇰
    🇵🇹
    
  

# 💰 ~1.6B Free Tokens / Month

> Stacking free tiers by hand is painful — dozens of SDKs, dozens of rate limits, and no idea how much you actually have. OmniRoute aggregates the **documented** free tiers of **40+ provider pools / 500+ models** into one honest number and shows it live on the dashboard (`/dashboard/free-tiers`).

- **~1.6B free tokens / month** (steady) — and **up to ~2.1B in your first month** with signup credits.
- **Pool-deduped, honest** — we count each shared free pool **once**, so the headline isn't inflated by rate-limit ceilings the way multi-billion competitor claims are. (Counting every rate limit 24/7 would read ~10B; we don't publish that.)
- **Plus the un-countable** — permanently-free, no-token-cap providers (SiliconFlow, Z.AI GLM-Flash, Kilo, OpenCode Zen…) and a **$10 OpenRouter top-up** that unlocks **+24M/mo**, both surfaced separately so they never inflate the headline.
- **Per-model breakdown**, **live used / remaining** for the current month, and a transparent **terms flag** per provider.

> Preview mockup — a real screenshot lands once the `/dashboard/free-tiers` page is validated. Full methodology (pool dedupe, credit tiers, provider terms): **[docs/reference/FREE_TIERS.md](docs/reference/FREE_TIERS.md)**.

# 💥 The Promise

> One endpoint. **237 providers.** Never stop building — and let OmniRoute pick the cheapest one that works.

  
    🚫 Never hit limitsAuto-fallback across 237 providers in milliseconds. Quota out? Next provider takes over — zero downtime.
    💸 Save up to 95% tokensRTK + Caveman stacked compression cuts 15–95% of eligible tokens (~89% avg on tool-heavy sessions).
    🆓 $0 to start50+ providers with a free tier, 11 free forever (Kiro, Qoder, Pollinations, LongCat…). No card needed.
  
  
    🔌 Every tool works16+ coding agents — Claude Code, Codex, Cursor, Cline, Copilot, Antigravity — through one config.
    🧩 One endpointOpenAI ↔ Claude ↔ Gemini ↔ Responses API translation. Point any tool at /v1 and it just works.
    🛡️ Production-gradeCircuit breakers, TLS stealth, MCP (87 tools), A2A, memory, guardrails, evals. 14,965 tests.
  

# 🤔 Why OmniRoute?

> Stop juggling 10 dashboards, dead API keys, and surprise bills.

| ❌ The daily pain                                      | ✅ How OmniRoute fixes it                                                     |
| ------------------------------------------------------ | ----------------------------------------------------------------------------- |
| 📉 Subscription quota expires unused every month       | **Maximize subscriptions** — track quota, use every token before reset        |
| 🛑 Rate limits stop you mid-coding                     | **4-tier auto-fallback** — Subscription → API → Cheap → Free, in milliseconds |
| 🔥 Tool outputs (`git diff`, `grep`, logs) burn tokens | **RTK + Caveman compression** — save 15–95% eligible tokens per request       |
| 💸 Expensive APIs ($20–50/mo per provider)             | **Cost-optimized routing** — auto-route to the cheapest viable model          |
| 🧰 Each AI tool wants its own setup                    | **One endpoint, every tool, one dashboard**                                   |
| 🌍 AI blocked in your country                          | **3-level proxy** + TLS fingerprint stealth — use AI from anywhere            |

```
┌──────────────────────────────────────────────────────────┐
│        Your IDE / CLI  (Claude Code, Cursor, Cline…)       │
└─────────────────────────┬──────────────────────────────────┘
                          │ http://localhost:20128/v1
                          ▼
┌──────────────────────────────────────────────────────────┐
│                  OmniRoute — Smart Router                  │
│  RTK + Caveman compression · 17 routing strategies         │
│  Circuit breakers · TLS stealth · MCP · A2A · Guardrails   │
└─────────────────────────┬──────────────────────────────────┘
        ┌─────────────┬────┴────────┬─────────────┐
        ▼ Tier 1      ▼ Tier 2      ▼ Tier 3       ▼ Tier 4
   SUBSCRIPTION     API KEY        CHEAP          FREE
   Claude Code,     DeepSeek,      GLM $0.5,      Kiro, Qoder,
   Codex, Copilot   Groq, xAI      MiniMax $0.2   Pollinations
   quota out? ───▶  budget hit? ─▶ budget hit? ─▶ always on
```

# 🎯 Combos — The Flagship

> A **combo** is a chain of models OmniRoute routes across **automatically**. Quota runs out, a provider fails, or costs spike — the combo silently slides to the next model. **This is what makes OmniRoute unbreakable.** 🛡️

### ⚡ Zero-config — just use `auto`

No combo to create. Set your model to `auto` (or a variant) and OmniRoute builds a virtual combo from your connected providers, scored live:

| Model ID       | What it optimizes for                                          |
| -------------- | -------------------------------------------------------------- |
| `auto`         | 🎯 Balanced default (LKGP — sticks to your last good provider) |
| `auto/coding`  | 🧑‍💻 Quality-first weights for code generation                   |
| `auto/fast`    | ⚡ Lowest latency first                                        |
| `auto/cheap`   | 💰 Cheapest per token first                                    |
| `auto/offline` | 🔋 Most quota / rate-limit headroom first                      |
| `auto/smart`   | 🔭 Quality-first + 10% exploration to discover better models   |

##

### 🔀 Or build your own — 17 routing strategies

| Goal                                    | Strategy / combo                                   |
| --------------------------------------- | -------------------------------------------------- |
| 🥇 Drain my subscription before paying  | `priority` / `fill-first`                          |
| ⚖️ Spread load across accounts          | `round-robin` · `weighted` · `p2c` · `least-used`  |
| 💸 Always cheapest viable model         | `cost-optimized` · `auto/cheap`                    |
| 🧠 Hand off long context between models | `context-relay` · `context-optimized`              |
| 🎲 Randomized / privacy routing         | `random` · `strict-random`                         |
| 🧬 Fan out to a panel + judge synthesis | `fusion`                                           |
| 📊 Route by remaining quota headroom    | `reset-window` · `headroom`                        |
| 🤖 Just make it smart                   | `auto` (9-factor scoring) · `lkgp` · `reset-aware` |

The Auto-Combo engine scores every candidate on **9 factors** (health, quota, cost, latency, success rate, freshness…) — see [`docs/routing/AUTO-COMBO.md`](docs/routing/AUTO-COMBO.md).

##

### 🧱 Resilience is built in (3 independent layers)

| Layer                      | Scope             | What it does                                                               |
| -------------------------- | ----------------- | -------------------------------------------------------------------------- |
| 🔌 **Circuit breaker**     | whole provider    | Stops hammering a provider that's failing upstream; auto-probes to recover |
| 💤 **Connection cooldown** | one account / key | Skips a rate-limited key while other keys keep serving                     |
| 🎯 **Model lockout**       | provider + model  | Quarantines just one quota-limited model, not the whole connection         |

```
Combo: "always-on"                         Strategy: priority
  1. cc/claude-opus-4-7   ← subscription (use it fully)
  2. cx/gpt-5.5           ← second subscription
  3. glm/glm-5.1          ← cheap backup ($0.5/1M)
  4. kr/claude-sonnet-4.5 ← FREE, unlimited (never fails)
Result: 4 layers of fallback = zero downtime
```

📖 [Auto-Combo Engine](docs/routing/AUTO-COMBO.md) · [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md)

# 🏆 What Sets OmniRoute Apart

| Feature                                | OmniRoute                                                           | Other routers |
| -------------------------------------- | ------------------------------------------------------------------- | ------------- |
| 🌐 Providers                           | **231**                                                             | 20–100        |
| 🆓 Free providers                      | **50+ (11 free forever)**                                           | 1–5           |
| 🔀 Routing strategies                  | **17** (priority, weighted, cost-optimized, context-relay, fusion…) | 1–3           |
| 🗜️ Token compression                   | **RTK + Caveman stacked (15–95%)**                                  | None / 20–40% |
| 🧰 Built-in MCP server                 | **87 tools, 3 transports, 30 scopes**                               | Rare          |
| 🤝 A2A agent protocol                  | **6 skills, JSON-RPC 2.0**                                          | None          |
| 🧠 Memory (FTS5 + vector)              | **Yes**                                                             | Rare          |
| 🛡️ Guardrails (PII, injection, vision) | **Yes**                                                             | Rare          |
| ☁️ Cloud agents                        | **Codex, Devin, Jules**                                             | None          |
| 🥷 TLS fingerprint stealth             | **JA3/JA4 via wreq-js**                                             | None          |
| 🖥️ Multi-platform                      | **Web · Desktop · Termux · PWA**                                    | Web only      |
| 🌍 i18n                                | **42 locales**                                                      | 0–4           |

📊 Detailed comparison vs LiteLLM, OpenRouter & Portkey → [`docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md`](docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md)

# ✨ What's New

> Recent highlights from **v3.8.20 → v3.8.39**. Full history in [`CHANGELOG.md`](CHANGELOG.md).

- **⚖️ Quota-Share routing** — a dedicated combo strategy that spreads load across accounts by _available quota_: Deficit-Round-Robin scheduling, per-connection `max_concurrent` with cooldown-wait queueing, multi-window usage buckets (5h / 7d / per-model), per-(key,model) caps, session stickiness for prompt-cache integrity, and proactive saturation from upstream token-usage headers. → [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md)
- **🤖 One-command CLI/agent setup** — a dedicated `setup-*` command configures each coding tool to route through OmniRoute (Claude Code, Codex, Cline, Continue, Cursor, Roo Code, Kilo Code, Crush, Goose, Qwen Code, Aider, OpenCode); `omniroute launch` / `omniroute launch-codex` are zero-config launchers. → [CLI Integrations](docs/guides/CLI-INTEGRATIONS.md)
- **🛰️ Remote mode** — drive a remote OmniRoute from any machine with scoped access tokens (`omniroute connect` / `omniroute contexts` / `omniroute tokens`), plus an `omniroute login antigravity` helper that runs Google "native/desktop" OAuth on your own machine and pastes a credential blob into a remote/VPS install (where the loopback redirect is unreachable). → [Remote Mode](docs/guides/REMOTE-MODE.md)
- **🧭 Smarter auto-routing** — OpenRouter-style `auto/:` combos (e.g. `auto/coding:fast`, `auto/reasoning:pro`), a **Fusion** strategy (fan out to a panel of models in parallel, then synthesize via a judge), **task-aware routing** (best-fit connection per task type), per-request `X-Route-Model` override, live Arena-ELO + models.dev model intelligence, per-step account allowlists, provider-wildcard combo steps, nested combo-ref execution, sticky weighted selection, and `web_search`-aware routing. → [Auto-Combo](docs/routing/AUTO-COMBO.md)
- **🗜️ Pluggable compression** — an async pipeline of **9 composable engines** with Compression Studios, an LLMLingua-2 ONNX engine and a heuristic/SLM two-tier **Ultra**, RTK, delegated Anthropic Context Editing, **Output Styles** (output-axis steering: terse-prose / less-code / terse-CJK), an **adaptive context-budget dial** (escalate only as far as needed to fit the context window), per-request `x-omniroute-compression` control, an opt-in offline eval harness, one-click **Headroom** proxy lifecycle management from the dashboard (Docker sidecar supported), a synthetic **compression playground** (Play lanes + A/B Compare with USD-capped fidelity verdicts), an opt-in **per-step fidelity gate** that rejects a lossy engine before it degrades the prompt, a **best-of-N candidate encoder** (GCF vs TOON — keep whichever is shorter, with an A/B bytes/token table in the studio), **CCR ranged/grep/stats retrieval** (pull an exact byte/line slice or summary of a stored block instead of re-expanding it), and a unified panel with named profiles + an active-profile selector. → [Compression](docs/compression/COMPRESSION_ENGINES.md)
- **🕵️ Transparent MITM decrypt (TPROXY)** — capture & translate traffic from CLIs that ignore proxy env vars, with a per-SNI certificate authority and a trust-store installer. → [MITM/TPROXY](docs/security/MITM-TPROXY-DECRYPT.md)
- **💸 Cost telemetry everywhere** — `X-OmniRoute-*` cost/usage headers on every endpoint (including media), a non-token cost engine, a cache-HIT `X-OmniRoute-Cost-Saved` header, and per-key USD spend quotas. → [API Reference](docs/reference/API_REFERENCE.md)
- **🧠 Memory you control** — opt-in int8 vector quantization (Qdrant + sqlite-vec), memory off by default, and a pe

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [diegosouzapw](https://github.com/diegosouzapw)
- **Source:** [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute)
- **License:** MIT
- **Homepage:** https://omniroute.online

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** yes
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** yes
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-diegosouzapw-omniroute
- Seller: https://agentstack.voostack.com/s/diegosouzapw
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
