All routes operationalMọi routes đang hoạt động

Unify every model behind one gateway.Hợp nhất mọi model sau một gateway.

One OpenAI-compatible API, CLI and MCP gateway in front of every provider — and every other router. Free models, $2 credit a month, and your own keys at no markup.Một API tương thích OpenAI, CLI và MCP gateway đứng trước mọi provider — và cả các router khác. Model miễn phí, $2 credit mỗi tháng, và key của bạn không bị cắt phí.

Free to start — $2 credit every month, no card. Unlock the shared pool for $1/mo or by donating one working key (earn up to 8% back).Miễn phí để bắt đầu — $2 credit mỗi tháng, không cần thẻ. Mở khóa pool chung với $1/tháng hoặc góp một key đang dùng được (nhận lại 5%).

Use it withDùng vớiOpenAIxAIAnthropic+ 169 models across 17 providers+ 169 model qua 17 provider
Reading asĐọc với tư cách
anyrouter ~ openai (python)
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://anyrouter.dev/api/v1",
    api_key=os.environ["ANYROUTER_API_KEY"],
)

resp = client.chat.completions.create(
    model="anthropic/claude-sonnet-4.6",
    messages=[{"role": "user", "content": "Hi"}],
)
ANYROUTER_API_KEYget keylấy key

Crowdsourced free tokens

We pool everyone's free & trial keys into one big shared pool — free tokens for everyone. Donate a key, grow the pool, ride free.

01

One endpoint. Every model.Một endpoint. Mọi model.

Keep your SDK. One upstream rate-limits — your request doesn't notice.Giữ nguyên SDK. Một upstream bị rate-limit — request của bạn không hề hay biết.

Point your SDK at one base URL and reference models as provider/model. Every call then runs the same pipeline: authenticate, pick an upstream, retry the next one on failure, cache what repeats, and log every attempt.Trỏ SDK tới một base URL và gọi model dạng provider/model. Mọi call chạy qua cùng một pipeline: xác thực, chọn upstream, thử upstream kế tiếp khi lỗi, cache phần lặp lại, và ghi log từng attempt.

Between your call and the model · per requestGiữa call và model · mỗi request

Key balancingCân bằng key

Spreads load across your keys and quarantines burned ones automatically — no client changes.Phân tải giữa các key và tự cách ly key bị burn — không sửa client.

Automatic failoverFailover tự động

Retries and falls back across providers and pooled keys the instant one errors or rate-limits.Thử lại và chuyển provider/key ngay khi lỗi hoặc rate-limit.

Prompt cachingPrompt caching

Reuses cached context across calls to cut repeat token cost and time-to-first-token.Tái sử dụng context đã cache để giảm chi phí token lặp và TTFT.

Per-request logsLog từng request

Every fallback hop, token count and cost captured per attempt — traced and queryable.Mọi hop fallback, token và chi phí theo từng attempt — truy vết và truy vấn được.

03

Everything in one gatewayMọi thứ trong một gateway

One key. Every capability you'd otherwise wire up yourself.Một key. Mọi khả năng bạn phải tự nối trước đây.

No add-ons, no tiers gating the basics. Every capability below ships on the same base URL and the same API key.Không add-on, không tier khóa tính năng cơ bản. Mọi khả năng dưới đây dùng chung base URL và API key.

04

Why teams switchVì sao team chuyển sang

Three things that are structurally hard to copy.Ba điều về cấu trúc khó sao chép.

Never taxedKhông bị cắt

Bring your own keys, keep every cent.Dùng key của bạn, giữ trọn từng xu.

Route through your own provider keys and AnyRouter takes no cut — no deposit fee, no per-token markup, no cap. What you pay upstream is exactly what you pay.Route qua key provider của bạn và AnyRouter không lấy phần — không phí nạp, không markup theo token, không cap. Trả upstream bao nhiêu thì trả đúng bấy nhiêu.

ObservableQuan sát được

Debug any request.Debug mọi request.

Every fallback hop, status, latency, token count and cost — captured per attempt. A “debug this request” trace nobody else ships at this tier.Mọi hop fallback, status, latency, token và chi phí — ghi theo từng attempt. Trace “debug request này” mà tier này ít ai ship.

PortableDi động

Your config follows you.Config đi theo bạn.

Keys, presets and skills live in AnyRouter, inject locally on demand, and wipe clean on exit. Same setup on your laptop, a server, CI, or a teammate's box.Keys, preset và skills sống trong AnyRouter, inject local khi cần, xóa sạch khi thoát. Cùng setup trên laptop, server, CI hay máy đồng đội.

05

Free models, funded by everyone

Shared key pool

Every member signs up for a provider's free/trial tier — NVIDIA NIM, the Gemini free tier, the Groq free tier, and more — and donates that key. The quotas add up into one pool that every member — you included — can call. Donors earn up to 8% back and their Start plan is free.

06

Unlock Start — two doors, same room

How to join

AnyRouter runs on Start $1/mo, or donate one free-tier provider key (donors ride free and earn up to 8% back). Every Start comes with $2 credit each month.

Create your account

Sign up in seconds — no card required. You land on the dashboard with your first API key ready to copy.

07

The cheapest way in

Three ways to start

A $2 monthly credit and free models come with Start — $1/mo, or free when you donate a provider key. Or bring your own keys at no markup, no card required.

Start plan

$2 credit / month

On Start — $1/mo or a donated key. Here's how far $2 goes on fast models.

Tokens for $2, blended 75% input / 25% output.

Start plan

Free models, 1000/day

Route via anyrouter/free on Start — $0 per token, up to 1000 requests/day.

North Mini Code
$0 in · $0 out
DeepSeek-R1
$0 in · $0 out
DeepSeek-V3
$0 in · $0 out
DeepSeek V3.1
$0 in · $0 out
DeepSeek V3.2
$0 in · $0 out
DeepSeek-V4-Flash
$0 in · $0 out
DeepSeek V4 Pro
$0 in · $0 out
DiffusionGemma 26B A4B
$0 in · $0 out
Gemma 4 26B A4B
$0 in · $0 out
Gemma-4-27B-it
$0 in · $0 out
Gemma 4 31B
$0 in · $0 out
Ling-3.0-flash
$0 in · $0 out
Bring your own keys

Your keys, no markup

11 providers ship a free tier — billed by the provider only, never by us.

OpenCode Zen
Cerebras
Google AI Studio
Mistral
NVIDIA NIM
Ollama Cloud
OpenRouter
QwenCloud
ZenMux
Hugging Face

When you need more

Full pricing
Start$1/mo$2 in credits every month and free models. $1/mo + payment fee, or unlock everything free by donating a provider key.
Prosold out$9/moFor developers running agents every day.
Maxsold out$45/moFor teams pushing volume on the biggest models.
08

Sign in with AnyRouter

Add AI to your app — your users bring their own AnyRouter

One OAuth button hands your app a scoped, temporary key per user. Every request is billed to that user's own account — no API keys for you to collect, store, or pay for.

Your app's sign-in screen

Click it — see what your users see.

  • OAuth 2.1 + PKCE

    Standard authorization-code flow, public client, no secret to leak. Works from a browser-only app.

  • Inference-only scoped key

    The token can run AI requests and read a basic profile — never keys, billing, or account settings.

  • Per-user billing & revocation

    Each user pays from their own credits or free tier, and can revoke your app any time from their dashboard.

09

169+ models · 17 providers · July 25, 2026

One growing catalog, always current

Anthropic's Claude Opus 5 — the flagship model for demanding reasoning, coding, and long-horizon agentic work — joins the catalog, BYOK-only like its Claude siblings.

1M

Anthropic's Claude Opus 5 is the flagship model for demanding reasoning, coding, and long-horizon agentic work — end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, and coordinating parallel subagents.

262KFreeMiễn phí

Ling-3.0-flash is inclusionAI's instant (instruct) model, a 124B-parameter Mixture-of-Experts model with roughly 5.1B activated parameters per token. It is designed with token efficiency and production-scale agentic inference as key priorities, delivering fast responses and strong execution across coding, document processing, and lightweight agent workflows. Served free via Novita's own $0 pricing, with BYOK as a fallback.

128KFreeMiễn phí

A 31 billion parameter model from NVIDIA served via NVIDIA NIM, tuned for efficient text generation and reasoning tasks.

1.0MZDR

Google's Gemini 3.5 Flash Lite is a cost-efficient, high-throughput multimodal model available as a partner-hosted SKU on Cloudflare Workers AI. Optimized for agentic workflows, complex coding, and long-horizon tasks across a 1M-token context window.

1.0MZDR

Google's Gemini 3.6 Flash is a frontier-level multimodal model available as a partner-hosted SKU on Cloudflare Workers AI. Combines fast inference with strong reasoning, coding, and multimodal understanding across a 1M-token context window.

1.0MFreeMiễn phí

A 118B total parameter Mixture-of-Experts coding model with 8B activated parameters per token, scoring 70.2% on Terminal-Bench 2.1 and 40.4% on DeepSWE. Supports tool calling, reasoning, thinking and no-thinking modes, with a context window of up to 1M tokens. Open-weight under the OpenMDW-1.1 License.

8KFreeMiễn phí

NVIDIA Llama-Nemotron-Embed-VL-1B-v2 is a vision-language embedding model for multimodal question-answering and retrieval over text, images, or combined image-text documents. A transformer encoder fine-tuned from Llama 3.2 1B with SigLip2 400M, it uses a tiling-based VLM architecture (Eagle 2 + nemoretriever-parse) for high-resolution image and complex visual-document understanding. Served via NVIDIA NIM.

33KFreeMiễn phí

NVIDIA Nemotron-3-Embed-1B is a 1.14B-parameter multilingual text embedding model (2048 dimensions, 34 languages) for semantic search, retrieval, and RAG. Built on Ministral-3-3B-Instruct; state-of-the-art on multilingual retrieval benchmarks (MMTEB 71.05, RTEB 72.38). Served via NVIDIA NIM.

1.0MFreeMiễn phí

Inkling is Thinking Machines Lab's first open-weights foundation model: a 975B-parameter multimodal Mixture-of-Experts transformer (41B active, 256 experts) with a 1M token context window, switchable reasoning effort, and tool use. Accepts text, image, and audio inputs; outputs text. Pretrained on 45T tokens. Served via NVIDIA NIM.

1.0M

Meta's Muse Spark 1.1 is a multimodal reasoning model built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context window. The model orchestrates multi-agent workflows — acting as a main agent that plans and delegates or as a subagent — and generalizes zero-shot to new tools, MCP servers, and custom skills. It supports structured output, parallel function calling, built-in search with citations, and configurable reasoning effort. Available with a user-owned OpenRouter API key (BYOK). NOTE: currently available to users in the United States only.

33KZDR

Moondream 3.1 is a fast, efficient 9B mixture-of-experts vision language model (2B active parameters) delivering frontier-level visual reasoning for object detection, pointing, OCR, and structured output. Served through the OpenAI chat-completions surface by translating image + prompt into the Moondream query skill.

1.0MFreeMiễn phí

Kimi K3 is an open-source multimodal reasoning model from Moonshot AI with a 1M-token context window and native visual understanding. It accepts text, image, and video inputs and is built for long-horizon software engineering, knowledge work, and deep reasoning, with particular strength where code meets visual and spatial reasoning — frontend, game development, and CAD workflows. Its architecture uses KDA and Attention Residuals for computational efficiency, and its thinking mode is always on.

131KFreeMiễn phí

NVIDIA Nemotron 3 Nano Omni 30B-A3B (reasoning) is a multimodal MoE model unifying video, audio, image, and text understanding for enterprise Q&A, summarization, transcription, OCR

1.1MZDR

OpenAI's GPT-5.6 Luna, a reasoning-focused model in the GPT-5.6 family optimized for cost-sensitive workloads, with a 1M+ token context window and tool-use support via the Responses API.

1.1M

OpenAI's GPT-5.6 Sol, a reasoning-focused model in the GPT-5.6 family with a 1M+ token context window, configurable reasoning effort, and tool-use support via the Responses API.

1.1M

OpenAI's GPT-5.6 Terra, a reasoning-focused model in the GPT-5.6 family that balances intelligence and cost, with a 1M+ token context window and tool-use support via the Responses API.

33KFreeMiễn phí

A lean, efficient text model from Mistral, optimized for low-latency generation with strong cache-hit performance.

33KFreeMiễn phí

A compact 30 billion parameter model from NVIDIA's Nemotron-3 family, tuned for fast, efficient text generation and understanding tasks.

262KFreeMiễn phí

The latest coding agent model in the 33B-A3B category from Poolside and a step forward from Laguna XS.2. Combines tool calling and reasoning with a compact footprint, offering a 256K context window and up to 32K output tokens. Quantized to FP8 for fast, cost-efficient agentic coding workflows. Subject to the OpenMDW-1.1 License.

500K

xAI's Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

64KFreeMiễn phí

DeepSeek's flagship reasoning model with chain-of-thought. Strong math, coding, and complex problem solving.

128KFreeMiễn phí

671B MoE general-purpose model with strong multilingual, coding, and reasoning performance.

128KFreeMiễn phí

Fast and cheap DeepSeek-V4 MoE variant via SiliconFlow free tier for high-throughput chat and coding.

128KFreeMiễn phí

Google Gemma 4 27B open-weight instruct model — strong general chat at low cost.

1M

Anthropic's Claude Opus 5 is the flagship model for demanding reasoning, coding, and long-horizon agentic work — end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, and coordinating parallel subagents.

262KFreeMiễn phí

Ling-3.0-flash is inclusionAI's instant (instruct) model, a 124B-parameter Mixture-of-Experts model with roughly 5.1B activated parameters per token. It is designed with token efficiency and production-scale agentic inference as key priorities, delivering fast responses and strong execution across coding, document processing, and lightweight agent workflows. Served free via Novita's own $0 pricing, with BYOK as a fallback.

128KFreeMiễn phí

A 31 billion parameter model from NVIDIA served via NVIDIA NIM, tuned for efficient text generation and reasoning tasks.

1.0MZDR

Google's Gemini 3.5 Flash Lite is a cost-efficient, high-throughput multimodal model available as a partner-hosted SKU on Cloudflare Workers AI. Optimized for agentic workflows, complex coding, and long-horizon tasks across a 1M-token context window.

1.0MZDR

Google's Gemini 3.6 Flash is a frontier-level multimodal model available as a partner-hosted SKU on Cloudflare Workers AI. Combines fast inference with strong reasoning, coding, and multimodal understanding across a 1M-token context window.

1.0MFreeMiễn phí

A 118B total parameter Mixture-of-Experts coding model with 8B activated parameters per token, scoring 70.2% on Terminal-Bench 2.1 and 40.4% on DeepSWE. Supports tool calling, reasoning, thinking and no-thinking modes, with a context window of up to 1M tokens. Open-weight under the OpenMDW-1.1 License.

8KFreeMiễn phí

NVIDIA Llama-Nemotron-Embed-VL-1B-v2 is a vision-language embedding model for multimodal question-answering and retrieval over text, images, or combined image-text documents. A transformer encoder fine-tuned from Llama 3.2 1B with SigLip2 400M, it uses a tiling-based VLM architecture (Eagle 2 + nemoretriever-parse) for high-resolution image and complex visual-document understanding. Served via NVIDIA NIM.

33KFreeMiễn phí

NVIDIA Nemotron-3-Embed-1B is a 1.14B-parameter multilingual text embedding model (2048 dimensions, 34 languages) for semantic search, retrieval, and RAG. Built on Ministral-3-3B-Instruct; state-of-the-art on multilingual retrieval benchmarks (MMTEB 71.05, RTEB 72.38). Served via NVIDIA NIM.

1.0MFreeMiễn phí

Inkling is Thinking Machines Lab's first open-weights foundation model: a 975B-parameter multimodal Mixture-of-Experts transformer (41B active, 256 experts) with a 1M token context window, switchable reasoning effort, and tool use. Accepts text, image, and audio inputs; outputs text. Pretrained on 45T tokens. Served via NVIDIA NIM.

1.0M

Meta's Muse Spark 1.1 is a multimodal reasoning model built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context window. The model orchestrates multi-agent workflows — acting as a main agent that plans and delegates or as a subagent — and generalizes zero-shot to new tools, MCP servers, and custom skills. It supports structured output, parallel function calling, built-in search with citations, and configurable reasoning effort. Available with a user-owned OpenRouter API key (BYOK). NOTE: currently available to users in the United States only.

33KZDR

Moondream 3.1 is a fast, efficient 9B mixture-of-experts vision language model (2B active parameters) delivering frontier-level visual reasoning for object detection, pointing, OCR, and structured output. Served through the OpenAI chat-completions surface by translating image + prompt into the Moondream query skill.

1.0MFreeMiễn phí

Kimi K3 is an open-source multimodal reasoning model from Moonshot AI with a 1M-token context window and native visual understanding. It accepts text, image, and video inputs and is built for long-horizon software engineering, knowledge work, and deep reasoning, with particular strength where code meets visual and spatial reasoning — frontend, game development, and CAD workflows. Its architecture uses KDA and Attention Residuals for computational efficiency, and its thinking mode is always on.

131KFreeMiễn phí

NVIDIA Nemotron 3 Nano Omni 30B-A3B (reasoning) is a multimodal MoE model unifying video, audio, image, and text understanding for enterprise Q&A, summarization, transcription, OCR

1.1MZDR

OpenAI's GPT-5.6 Luna, a reasoning-focused model in the GPT-5.6 family optimized for cost-sensitive workloads, with a 1M+ token context window and tool-use support via the Responses API.

1.1M

OpenAI's GPT-5.6 Sol, a reasoning-focused model in the GPT-5.6 family with a 1M+ token context window, configurable reasoning effort, and tool-use support via the Responses API.

1.1M

OpenAI's GPT-5.6 Terra, a reasoning-focused model in the GPT-5.6 family that balances intelligence and cost, with a 1M+ token context window and tool-use support via the Responses API.

33KFreeMiễn phí

A lean, efficient text model from Mistral, optimized for low-latency generation with strong cache-hit performance.

33KFreeMiễn phí

A compact 30 billion parameter model from NVIDIA's Nemotron-3 family, tuned for fast, efficient text generation and understanding tasks.

262KFreeMiễn phí

The latest coding agent model in the 33B-A3B category from Poolside and a step forward from Laguna XS.2. Combines tool calling and reasoning with a compact footprint, offering a 256K context window and up to 32K output tokens. Quantized to FP8 for fast, cost-efficient agentic coding workflows. Subject to the OpenMDW-1.1 License.

500K

xAI's Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

64KFreeMiễn phí

DeepSeek's flagship reasoning model with chain-of-thought. Strong math, coding, and complex problem solving.

128KFreeMiễn phí

671B MoE general-purpose model with strong multilingual, coding, and reasoning performance.

128KFreeMiễn phí

Fast and cheap DeepSeek-V4 MoE variant via SiliconFlow free tier for high-throughput chat and coding.

128KFreeMiễn phí

Google Gemma 4 27B open-weight instruct model — strong general chat at low cost.