AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

OmniRoute

mcp-diegosouzapw-omniroute · by diegosouzapw

Never stop coding. Free AI gateway: one endpoint, 160+ providers (50+ free), connect Claude Code, Codex, Cursor, Cline & Copilot to FREE Claude/GPT/Gemini. RTK+Caveman stacked compression saves 15-95% tokens, smart auto-fallback, MCP/A2A, multimodal APIs, Desktop/PWA.

No reviews yet
0 installs
43 views
0.0% view→install

Install

$ agentstack add mcp-diegosouzapw-omniroute

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-diegosouzapw-omniroute)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of OmniRoute? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

🚀 OmniRoute — The Free AI Gateway

Never stop coding. Connect every AI tool to 237 providers50+ free — through one endpoint.

Plug Claude Code, Codex, Cursor, Cline, Copilot & Antigravity into FREE Claude / GPT / Gemini. Auto-fallback.

RTK + Caveman compression saves 15–95% tokens. Never hit limits.

~1.6B documented free tokens/month — up to ~2.1B in your first month with signup credits — aggregated across the free tiers, plus a long tail of permanently-free, no-cap providers, and the compression above stretches every one further. ([how we count →](docs/reference/FREE_TIERS.md#tldr--how-much-free-inference-does-omniroute-actually-aggregate))

[](#-231-ai-providers--50-free) [](#-231-ai-providers--50-free) [](docs/reference/FREE_TIERS.md) [](#%EF%B8%8F-save-1595-tokens--automatically) [](#-combos--the-flagship) [](#-quick-start)

💬 Join the community

[](https://discord.gg/EkzRkpzKYt) [](https://t.me/omnirouteOficial) [](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t) [](https://chat.whatsapp.com/BTGJXIyjeNIIgExvTMGGhI)

Questions, provider tips, roadmap & support → Discord · Telegram · WhatsApp 🌍 Global / 🇧🇷 Brasil

[](https://www.npmjs.com/package/omniroute) [](LICENSE) [](package.json) [](https://github.com/diegosouzapw/OmniRoute)

[](https://www.npmjs.com/package/omniroute)

[](https://hub.docker.com/r/diegosouzapw/omniroute)

[](https://omniroute.online)

[🚀 Quick Start](#-quick-start) • [🎯 Combos](#-combos--the-flagship) • [🌐 Providers](#-231-ai-providers--50-free) • [🔌 CLI & MCP](#-full-cli--a2a--mcp) • [🗜️ Compression](#%EF%B8%8F-save-1595-tokens--automatically) • 🌍 Website

[💥 The Promise](#-the-promise) • [🤔 Why](#-why-omniroute) • [🏆 What Sets Apart](#-what-sets-omniroute-apart) • [🤖 Compatible CLIs](#-compatible-clis--coding-agents) • [🖥️ Where It Runs](#%EF%B8%8F-where-omniroute-runs--anywhere) • [🔒 Private](#-private--local-first) • [🎬 In Action](#-omniroute-in-action) • [📚 Explore More](#-explore-more) • [📧 Support](#-support--community)

🌐 Available in 41+ languages

🇺🇸 🇧🇷 🇪🇸 🇫🇷 🇮🇹 🇷🇺 🇨🇳 🇹🇼 🇩🇪 🇯🇵 🇰🇷 🇮🇳

🇹🇭 🇻🇳 🇮🇩 🇲🇾 🇵🇭 🇸🇦 🇮🇱 🇦🇿 🇺🇦 🇵🇱 🇨🇿

🇳🇱 🇧🇬 🇩🇰 🇫🇮 🇳🇴 🇸🇪 🇭🇺 🇷🇴 🇸🇰 🇵🇹

💰 ~1.6B Free Tokens / Month

> Stacking free tiers by hand is painful — dozens of SDKs, dozens of rate limits, and no idea how much you actually have. OmniRoute aggregates the documented free tiers of 40+ provider pools / 500+ models into one honest number and shows it live on the dashboard (/dashboard/free-tiers).

  • ~1.6B free tokens / month (steady) — and up to ~2.1B in your first month with signup credits.
  • Pool-deduped, honest — we count each shared free pool once, so the headline isn't inflated by rate-limit ceilings the way multi-billion competitor claims are. (Counting every rate limit 24/7 would read ~10B; we don't publish that.)
  • Plus the un-countable — permanently-free, no-token-cap providers (SiliconFlow, Z.AI GLM-Flash, Kilo, OpenCode Zen…) and a $10 OpenRouter top-up that unlocks +24M/mo, both surfaced separately so they never inflate the headline.
  • Per-model breakdown, live used / remaining for the current month, and a transparent terms flag per provider.

> Preview mockup — a real screenshot lands once the /dashboard/free-tiers page is validated. Full methodology (pool dedupe, credit tiers, provider terms): [docs/reference/FREETIERS.md](docs/reference/FREETIERS.md).

💥 The Promise

> One endpoint. 237 providers. Never stop building — and let OmniRoute pick the cheapest one that works.

🚫 Never hit limitsAuto-fallback across 237 providers in milliseconds. Quota out? Next provider takes over — zero downtime. 💸 Save up to 95% tokensRTK + Caveman stacked compression cuts 15–95% of eligible tokens (~89% avg on tool-heavy sessions). 🆓 $0 to start50+ providers with a free tier, 11 free forever (Kiro, Qoder, Pollinations, LongCat…). No card needed.

🔌 Every tool works16+ coding agents — Claude Code, Codex, Cursor, Cline, Copilot, Antigravity — through one config. 🧩 One endpointOpenAI ↔ Claude ↔ Gemini ↔ Responses API translation. Point any tool at /v1 and it just works. 🛡️ Production-gradeCircuit breakers, TLS stealth, MCP (87 tools), A2A, memory, guardrails, evals. 14,965 tests.

🤔 Why OmniRoute?

> Stop juggling 10 dashboards, dead API keys, and surprise bills.

| ❌ The daily pain | ✅ How OmniRoute fixes it | | ------------------------------------------------------ | ----------------------------------------------------------------------------- | | 📉 Subscription quota expires unused every month | Maximize subscriptions — track quota, use every token before reset | | 🛑 Rate limits stop you mid-coding | 4-tier auto-fallback — Subscription → API → Cheap → Free, in milliseconds | | 🔥 Tool outputs (git diff, grep, logs) burn tokens | RTK + Caveman compression — save 15–95% eligible tokens per request | | 💸 Expensive APIs ($20–50/mo per provider) | Cost-optimized routing — auto-route to the cheapest viable model | | 🧰 Each AI tool wants its own setup | One endpoint, every tool, one dashboard | | 🌍 AI blocked in your country | 3-level proxy + TLS fingerprint stealth — use AI from anywhere |

┌──────────────────────────────────────────────────────────┐
│        Your IDE / CLI  (Claude Code, Cursor, Cline…)       │
└─────────────────────────┬──────────────────────────────────┘
                          │ http://localhost:20128/v1
                          ▼
┌──────────────────────────────────────────────────────────┐
│                  OmniRoute — Smart Router                  │
│  RTK + Caveman compression · 17 routing strategies         │
│  Circuit breakers · TLS stealth · MCP · A2A · Guardrails   │
└─────────────────────────┬──────────────────────────────────┘
        ┌─────────────┬────┴────────┬─────────────┐
        ▼ Tier 1      ▼ Tier 2      ▼ Tier 3       ▼ Tier 4
   SUBSCRIPTION     API KEY        CHEAP          FREE
   Claude Code,     DeepSeek,      GLM $0.5,      Kiro, Qoder,
   Codex, Copilot   Groq, xAI      MiniMax $0.2   Pollinations
   quota out? ───▶  budget hit? ─▶ budget hit? ─▶ always on

🎯 Combos — The Flagship

> A combo is a chain of models OmniRoute routes across automatically. Quota runs out, a provider fails, or costs spike — the combo silently slides to the next model. This is what makes OmniRoute unbreakable. 🛡️

⚡ Zero-config — just use auto

No combo to create. Set your model to auto (or a variant) and OmniRoute builds a virtual combo from your connected providers, scored live:

| Model ID | What it optimizes for | | -------------- | -------------------------------------------------------------- | | auto | 🎯 Balanced default (LKGP — sticks to your last good provider) | | auto/coding | 🧑‍💻 Quality-first weights for code generation | | auto/fast | ⚡ Lowest latency first | | auto/cheap | 💰 Cheapest per token first | | auto/offline | 🔋 Most quota / rate-limit headroom first | | auto/smart | 🔭 Quality-first + 10% exploration to discover better models |

##

🔀 Or build your own — 17 routing strategies

| Goal | Strategy / combo | | --------------------------------------- | -------------------------------------------------- | | 🥇 Drain my subscription before paying | priority / fill-first | | ⚖️ Spread load across accounts | round-robin · weighted · p2c · least-used | | 💸 Always cheapest viable model | cost-optimized · auto/cheap | | 🧠 Hand off long context between models | context-relay · context-optimized | | 🎲 Randomized / privacy routing | random · strict-random | | 🧬 Fan out to a panel + judge synthesis | fusion | | 📊 Route by remaining quota headroom | reset-window · headroom | | 🤖 Just make it smart | auto (9-factor scoring) · lkgp · reset-aware |

The Auto-Combo engine scores every candidate on 9 factors (health, quota, cost, latency, success rate, freshness…) — see [docs/routing/AUTO-COMBO.md](docs/routing/AUTO-COMBO.md).

##

🧱 Resilience is built in (3 independent layers)

| Layer | Scope | What it does | | -------------------------- | ----------------- | -------------------------------------------------------------------------- | | 🔌 Circuit breaker | whole provider | Stops hammering a provider that's failing upstream; auto-probes to recover | | 💤 Connection cooldown | one account / key | Skips a rate-limited key while other keys keep serving | | 🎯 Model lockout | provider + model | Quarantines just one quota-limited model, not the whole connection |

Combo: "always-on"                         Strategy: priority
  1. cc/claude-opus-4-7   ← subscription (use it fully)
  2. cx/gpt-5.5           ← second subscription
  3. glm/glm-5.1          ← cheap backup ($0.5/1M)
  4. kr/claude-sonnet-4.5 ← FREE, unlimited (never fails)
Result: 4 layers of fallback = zero downtime

📖 [Auto-Combo Engine](docs/routing/AUTO-COMBO.md) · [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md)

🏆 What Sets OmniRoute Apart

| Feature | OmniRoute | Other routers | | -------------------------------------- | ------------------------------------------------------------------- | ------------- | | 🌐 Providers | 231 | 20–100 | | 🆓 Free providers | 50+ (11 free forever) | 1–5 | | 🔀 Routing strategies | 17 (priority, weighted, cost-optimized, context-relay, fusion…) | 1–3 | | 🗜️ Token compression | RTK + Caveman stacked (15–95%) | None / 20–40% | | 🧰 Built-in MCP server | 87 tools, 3 transports, 30 scopes | Rare | | 🤝 A2A agent protocol | 6 skills, JSON-RPC 2.0 | None | | 🧠 Memory (FTS5 + vector) | Yes | Rare | | 🛡️ Guardrails (PII, injection, vision) | Yes | Rare | | ☁️ Cloud agents | Codex, Devin, Jules | None | | 🥷 TLS fingerprint stealth | JA3/JA4 via wreq-js | None | | 🖥️ Multi-platform | Web · Desktop · Termux · PWA | Web only | | 🌍 i18n | 42 locales | 0–4 |

📊 Detailed comparison vs LiteLLM, OpenRouter & Portkey → [docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md](docs/comparison/OMNIROUTEVSALTERNATIVES.md)

✨ What's New

> Recent highlights from v3.8.20 → v3.8.39. Full history in [CHANGELOG.md](CHANGELOG.md).

  • ⚖️ Quota-Share routing — a dedicated combo strategy that spreads load across accounts by available quota: Deficit-Round-Robin scheduling, per-connection max_concurrent with cooldown-wait queueing, multi-window usage buckets (5h / 7d / per-model), per-(key,model) caps, session stickiness for prompt-cache integrity, and proactive saturation from upstream token-usage headers. → [Resilience Guide](docs/architecture/RESILIENCE_GUIDE.md)
  • 🤖 One-command CLI/agent setup — a dedicated setup-* command configures each coding tool to route through OmniRoute (Claude Code, Codex, Cline, Continue, Cursor, Roo Code, Kilo Code, Crush, Goose, Qwen Code, Aider, OpenCode); omniroute launch / omniroute launch-codex are zero-config launchers. → [CLI Integrations](docs/guides/CLI-INTEGRATIONS.md)
  • 🛰️ Remote mode — drive a remote OmniRoute from any machine with scoped access tokens (omniroute connect / omniroute contexts / omniroute tokens), plus an omniroute login antigravity helper that runs Google "native/desktop" OAuth on your own machine and pastes a credential blob into a remote/VPS install (where the loopback redirect is unreachable). → [Remote Mode](docs/guides/REMOTE-MODE.md)
  • 🧭 Smarter auto-routing — OpenRouter-style auto/: combos (e.g. auto/coding:fast, auto/reasoning:pro), a Fusion strategy (fan out to a panel of models in parallel, then synthesize via a judge), task-aware routing (best-fit connection per task type), per-request X-Route-Model override, live Arena-ELO + models.dev model intelligence, per-step account allowlists, provider-wildcard combo steps, nested combo-ref execution, sticky weighted selection, and web_search-aware routing. → [Auto-Combo](docs/routing/AUTO-COMBO.md)
  • 🗜️ Pluggable compression — an async pipeline of 9 composable engines with Compression Studios, an LLMLingua-2 ONNX engine and a heuristic/SLM two-tier Ultra, RTK, delegated Anthropic Context Editing, Output Styles (output-axis steering: terse-prose / less-code / terse-CJK), an adaptive context-budget dial (escalate only as far as needed to fit the context window), per-request x-omniroute-compression control, an opt-in offline eval harness, one-click Headroom proxy lifecycle management from the dashboard (Docker sidecar supported), a synthetic compression playground (Play lanes + A/B Compare with USD-capped fidelity verdicts), an opt-in per-step fidelity gate that rejects a lossy engine before it degrades the prompt, a best-of-N candidate encoder (GCF vs TOON — keep whichever is shorter, with an A/B bytes/token table in the studio), CCR ranged/grep/stats retrieval (pull an exact byte/line slice or summary of a stored block instead of re-expanding it), and a unified panel with named profiles + an active-profile selector. → [Compression](docs/compression/COMPRESSION_ENGINES.md)
  • 🕵️ Transparent MITM decrypt (TPROXY) — capture & translate traffic from CLIs that ignore proxy env vars, with a per-SNI certificate authority and a trust-store installer. → [MITM/TPROXY](docs/security/MITM-TPROXY-DECRYPT.md)
  • 💸 Cost telemetry everywhereX-OmniRoute-* cost/usage headers on every endpoint (including media), a non-token cost engine, a cache-HIT X-OmniRoute-Cost-Saved header, and per-key USD spend quotas. → [API Reference](docs/reference/API_REFERENCE.md)
  • 🧠 Memory you control — opt-in int8 vector quantization (Qdrant + sqlite-vec), memory off by default, and a pe

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.