Keep requests moving
When a target is unavailable or out of quota, OmniRoute selects the next healthy target.
Free & open-source AI gateway
Route across 352 providers through one OpenAI-compatible endpoint, with automatic fallback.
OmniRoute × Cheaper Inference
Combine OmniRoute’s automatic fallback and developer tools with Cheaper Inference’s cost-ranked marketplace. Every eligible request starts with the cheapest available provider route.
| Model | Input price | Output price | Savings |
|---|---|---|---|
| gpt-5.6-lunaOpenAI · text | 60% off | ||
| gpt-5.6-solOpenAI · text | 50% off | ||
| glm-5.3Z.ai · text | 45% off | ||
| deepseek-v4-flashDeepSeek · text | 40% off | ||
| deepseek-v4-flash-0731DeepSeek · text | 40% off | ||
| glm-5.3-flashZ.ai · text | 30% off |
Core controls
When a target is unavailable or out of quota, OmniRoute selects the next healthy target.
A 12-engine pipeline reduces eligible context by 15–95%. The documented tool-heavy example reduced it by ~89%.
90+ free tiers and 56 recurring or keyless free-forever providers. No credit card needed for the keyless quick start.
Configure Claude Code, Codex, Cursor, Cline, Copilot, Antigravity and other supported clients through one endpoint.
Accept OpenAI, Claude, Gemini and Responses API requests at `/v1`.
Circuit breakers, encrypted credentials, MCP, A2A, memory, guardrails and evals ship with the self-hosted gateway.
The catalog
A catalog spanning chat, media, search, local runtimes, cloud agents and system integrations.
352 providers+
Quick start
Install the package, launch the gateway, then point any OpenAI-compatible client at port 20128.
Install the `omniroute` npm package globally.
$ npm install -g omnirouteadded 1 package
Launch the API and dashboard together on port 20128.
$ omniroute▸ dashboard ✓ http://localhost:20128/dashboard▸ api ...... ✓ serving on :20128
Use `http://localhost:20128/v1`; the keyless `auto` route works on a fresh install.
$ curl localhost:20128/v1/models✓ models listed
Gateway controls
Configure fallback, quota, compression, memory, MCP and A2A from the local dashboard.
Set a model to `auto` or build your own combo with 19 strategies and tiered fallback.
Use a ready-made mode or combine 19 strategies around the goal that matters.
Persistent conversational memory with FTS5 keyword and Qdrant vector recall.
Provider breaker, connection cooldown and model lockout isolate failures at the smallest useful scope.
Pool-deduped accounting across 90+ free tiers.
Published free-tier capacity, deduplicated.
12 composable engines for tool output, context and response style.
Reduce eligible tokens by 15–95%.
A built-in server exposing 110 gateway tools across 33 scopes.
110 tools · 33 scopes · stdio / HTTP
A JSON-RPC agent server with 6 skills.
6 skills · JSON-RPC 2.0 · Agent Card
Client compatibility
Point an OpenAI-compatible CLI or editor at one local base URL; OmniRoute handles the providers behind it.
http://localhost:20128/v1
curl http://localhost:20128/v1/chat/completions \
-H "Authorization: Bearer $OMNIROUTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"auto","messages":[{"role":"user","content":"Hello"}]}'OmniCopilot
OmniCopilot adds OmniRoute models to the existing model picker.
Works in VS Code stable and Insiders. Install from the Marketplace or Extensions view.
For Cursor, Windsurf, VSCodium, Theia, code-server, Gitpod and other compatible editors.
Provider models can run through OmniRoute in Copilot Chat without routing inference through a Copilot plan.
Common failure modes
See which control addresses each common failure mode.
Documentation snapshot
Capabilities checked against the linked project documentation on 2026-08-30.
| Capability | OmniRoute (opens in a new tab) | 9Router (opens in a new tab) | LiteLLM (opens in a new tab) | CLIProxyAPI (opens in a new tab) |
|---|---|---|---|---|
| Providers | 352 | 40+ | 100+ | 8+ upstreams |
| Routing strategies | 19 | 3-tier | retry / priority | round-robin / fill-first |
| Tiered fallback with UI | Yes | Yes | Manual policy | Multi-account |
| Token compression | 12 engines · 15–95% | RTK / Headroom / Caveman | No built-in pipeline | HTTP compression |
| Capability | OmniRoute (opens in a new tab) | 9Router (opens in a new tab) | LiteLLM (opens in a new tab) | CLIProxyAPI (opens in a new tab) |
|---|---|---|---|---|
| Built-in MCP server | 110 gateway tools | No | MCP client / gateway | No |
| A2A protocol | Server · 6 skills | No | A2A client / gateway | No |
| Cloud agents | Codex / Cursor / Devin / Jules | No | No | Sidecar integrations |
| Capability | OmniRoute (opens in a new tab) | 9Router (opens in a new tab) | LiteLLM (opens in a new tab) | CLIProxyAPI (opens in a new tab) |
|---|---|---|---|---|
| Failure-isolation layers | 3: provider / connection / model | Cooldown | Cooldown / retry | Cooldown |
| Persistent memory | FTS5 + vector | No | Semantic cache | No |
| Guardrails | PII / injection / vision | No | Yes | Cloaking controls |
| Eval framework | Built in | No | No | No |
| TLS fingerprint controls | wreq-js | No | No | uTLS |
| Capability | OmniRoute (opens in a new tab) | 9Router (opens in a new tab) | LiteLLM (opens in a new tab) | CLIProxyAPI (opens in a new tab) |
|---|---|---|---|---|
| Dashboard stack | Next.js 16 | Next.js 16 | Next.js | Next.js / React |
| UI locales | 42 | 10 | Not documented | 3 |
| OAuth providers | 25 | 6 | SSO | CLI auth |
| Self-hostable | Yes | Yes | Yes | Yes |
| License | MIT | MIT | MIT + enterprise | MIT |
Public README and official documentation snapshot; verify linked sources before purchase or deployment decisions.
Deployment
Install it on a laptop, server or phone, or connect through a supported editor and plugin.
One command on any supported desktop OS
npm install -g omnirouteOpen guide (opens in a new tab)Multi-arch AMD64 and ARM64 image
docker run … diegosouzapw/omnirouteOpen guide (opens in a new tab)Native window and system tray on Windows, macOS and Linux
npm run electron:buildOpen guide (opens in a new tab)Raspberry Pi, ARM servers and Apple Silicon
native arm64Runs on Android without root
pkg install nodejs && npx -y omnirouteOpen guide (opens in a new tab)Fullscreen, offline-ready and installable from a browser
Add to Home ScreenOpen guide (opens in a new tab)Native OpenCode provider plugin
@omniroute/opencode-providerOpen guide (opens in a new tab)OmniRoute models in the Copilot Chat picker
Install OmniCopilotOpen guide (opens in a new tab)Hack on it and contribute
npm install && npm run devOpen guide (opens in a new tab)Routing
Use a ready-made mode or combine 19 strategies around the goal that matters.
autoBalanced defaultcurl http://localhost:20128/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"auto","messages":[{"role":"user","content":"Write a TypeScript retry helper."}]}'Routing alias auto selected. Objective: Balanced default. Fallback enabled.
Use quota you already have before touching paid APIs.
priority · fill-firstBalance traffic across keys and accounts to reduce rate-limit pressure.
weighted · round-robin · p2c · least-used · headroomRoute to the lowest-cost eligible target.
cost-optimizedKeep long contexts on models that fit, then relay when needed.
context-relay · context-optimized · cache-optimizedDistribute eligible traffic across targets.
random · strict-randomUse 15-factor scoring, last-known-good and reset-aware windows.
reset-window · reset-aware · lkgp · auto · fusion · pipelineFree-tier accounting
455 catalog entries are deduplicated into 40 recurring pools; only published positive budgets feed the headline.
1.51B
Pool-deduped accounting keeps shared quotas from being counted more than once. No-cap providers stay visible without inflating the token total. Checked 2026-08-30.
View the calculation (opens in a new tab)Context compression
Tool output, repeated context and bloated structured data can be compressed before they reach the model.
The function should iterate over the list of items, and for each item, it should check whether the value is greater than zero, and if so, add it to the running total.sum all list items where value > 0x-omniroute-compression: standardStandard compression selected. Published range ~30%.
Control precedence: request header → combo → profile → adaptive policy → default.
Failure isolation
A bad key should not kill a provider. A limited model should not kill a connection.
Network controls
Proxy scopes control network routing while local-first storage keeps operational data on your machine.
Route globally, per provider or per individual connection.
Match supported provider clients down to selected transport characteristics.
Keys, usage and history stay on the host you control.
Local dashboard
Manage providers, combos, analytics, health and costs at `localhost:20128/dashboard`.
Agent protocols
OmniRoute can expose its gateway to MCP clients, A2A networks and supported cloud coding agents.
110 tools across 33 scopes over stdio and HTTP transports.
110 tools · 33 scopes · stdio / HTTP
claude mcp add omniroute --type http --url http://localhost:20128/api/mcp/stream6 skills for smart routing, quota, discovery, cost analysis and health reporting over JSON-RPC 2.0.
6 skills · JSON-RPC 2.0 · Agent Card
Connect supported Codex, Cursor, Devin and Jules workflows through one orchestration surface.
Codex · Cursor · Devin · Jules
Additional services
Run Bifrost, 9Router or CLIProxy as managed sidecars, or operate a full cluster profile.
Connect Notion and Obsidian as MCP-backed context sources.
Prompt-injection checks, opt-in PII redaction and normalized safety filters without provider lock-in.
Track streaks, achievements and live savings alongside operational analytics.
Guides
Walkthroughs cover installation, free-provider setup and connecting an IDE.
Built in the open
See stars, forks, contributors and releases from the checked-in GitHub snapshot.
Snapshot from 2026-09-01
Chat, support and roadmap
(opens in a new tab)News and announcements
(opens in a new tab)Global community
(opens in a new tab)Comunidade Brasil
(opens in a new tab)OmniRoute by Cheaper Inference
Self-host OmniRoute, or start with Cheaper Inference for a hosted cost-ranked gateway.