Zilmac Blog
← Back to Tech Practice

What Is OmniRoute? Complete Guide to One API for 290+ AI Models (2026)

API & Agents ·~14 min read

Self-hosted AI gateway dashboard and multi-model API routing topology for OmniRoute unified provider access

Change one base_url in Claude Code, Cursor, or your own Agent and you can switch between Claude, GPT, Gemini, DeepSeek, Kimi, and hundreds of other models on a single OpenAI-compatible API—that is not exclusive to OpenRouter. OmniRoute is among the fastest-growing open-source AI gateways on GitHub in 2026: MIT-licensed, self-hosted, with a catalog of 290+ providers and 500+ models (90+ with a free tier). This guide covers what it is, your first successful API call, Claude Code wiring, and a production security checklist.

290+
Configurable AI provider catalog
500+
Chat, embedding, image, and more
:20128
Default dashboard & API port

What OmniRoute is and what it solves

In short: OmniRoute is an LLM reverse proxy on your own machine. It exposes unified /v1/chat/completions (plus Anthropic /v1/messages, Gemini /v1beta/models, and more) upstream, and connects to your provider API keys, OAuth subscriptions, or free pools downstream.

Typical use cases

  • Multi-key fallback when a Claude subscription hits limits
  • One Agent entry for Claude Code, Cursor, and Cline with shared audit logs
  • Data stays inside—no third-party proxy in the request path
  • Budgets & quotas—see our Auto Combo budget routing guide
API request logs and multi-model routing on a developer laptop
Self-hosted gateway value: see which provider handled each request locally

OmniRoute vs. OpenRouter vs. LiteLLM

DimensionOmniRouteOpenRouterLiteLLM Proxy
HostingSelf-hostedCloud SaaSSelf-hosted (Python)
BillingProvider rates + your infraPrepaid tokens + ~5.5% top-upProvider rates + your infra
Catalog290+ providers / 500+ models300+ hosted modelsWhat you configure
Best forLocal gateway, mixed subscriptions, budgetsFast experiments, no opsMinimal Python proxy

For per-token price tables, read our OpenRouter pricing comparison. For Agent request budgets on OmniRoute, see Auto Combo budget routing.

Install & start (Docker / npm)

Docker (recommended)

docker run -d --name omniroute \
  -p 20128:20128 \
  -v omniroute-data:/data \
  diegosouzapw/omniroute

Open http://localhost:20128 → /dashboard.

npm (quick local trial)

npm install -g omniroute
omniroute

Cloud Mac tip

Run OmniRoute on a Zilmac cloud Mac for a stable IP and SSH access—pair LLM gateway with xcodebuild on the same or a sibling host.

Dashboard in four steps

  1. Create an API Key—all clients use Authorization: Bearer <key>.
  2. Connect providers—OAuth, paste API keys, or enable free no-auth pools (trial only).
  3. Point clients at http://<host>:20128/v1.
  4. Monitor usage—management routes without a key must return 401.

List all models: GET /v1/models

export OMNI_KEY="your-omniroute-api-key"
curl -s http://localhost:20128/v1/models \
  -H "Authorization: Bearer $OMNI_KEY" | jq '.data | length'

First API call (curl / Python / Node)

Use provider/model IDs. Missing prefix may be auto-added; mismatches return 400.

curl

curl -s http://localhost:20128/v1/chat/completions \
  -H "Authorization: Bearer $OMNI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"anthropic/claude-sonnet-4-6","messages":[{"role":"user","content":"Explain OmniRoute in one sentence"}],"max_tokens":256}'

Python

from openai import OpenAI
client = OpenAI(api_key="your-key", base_url="http://localhost:20128/v1")
print(client.chat.completions.create(
    model="openai/gpt-5.4",
    messages=[{"role": "user", "content": "Hello"}],
).choices[0].message.content)

Streaming

curl -N http://localhost:20128/v1/chat/completions \
  -H "Authorization: Bearer $OMNI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-chat","messages":[{"role":"user","content":"Count to 5"}],"stream":true}'

Anthropic Messages

Native Anthropic clients can use POST /v1/messages. Full endpoint list: API Reference.

auto/* intent aliases & fallback

Aliases like auto/best-coding pick a healthy model under quota—no client change when you swap providers. Use budget routing in production to avoid unwanted free-tier fallback.

curl -s http://localhost:20128/v1/chat/completions \
  -H "Authorization: Bearer $OMNI_KEY" -H "Content-Type: application/json" \
  -d '{"model":"auto/best-coding","messages":[{"role":"user","content":"quicksort"}]}'

Claude Code, Cursor & Copilot

ClientConfigNotes
Claude CodeANTHROPIC_BASE_URL=http://host:20128OAuth or API routing
CursorOverride OpenAI Base URL → /v1OmniRoute key as API key
Cline / ContinueOpenAI-compatible providerTry auto/best-coding

Production & security checklist

  • Upgrade to latest release (v3.8.50+); remove default passwords;
  • Never expose port 20128 on the public internet without TLS;
  • Per-member API keys, token limits, IP allowlists;
  • Treat free pools as fallback only; backup /data volume.

FAQ

400 model not found?

Check provider/model spelling and provider connection in the dashboard. Try auto/best-free to validate the chain.

Ollama locally?

Yes—POST /v1/api/chat integrates local models into the same routing table.

Use with OpenRouter?

Yes—OmniRoute can treat OpenRouter as a downstream provider for a hybrid setup.

Gateway for models, cloud Mac for builds

OmniRoute picks the LLM; iOS/macOS pipelines still need Xcode. Zilmac cloud Mac handles signing, notarization, and TestFlight alongside Agent workflows.

View cloud Mac plans

Limited Offer

Zilmac

Cloud Mac, remote dev, and Mac VPS—Apple toolchain support for iOS and cross-platform teams.

Back to Home
Limited Offer View Plans