Choose a model to open its comparison page.
Free API pricing covers 17 API providers, from $0.0010/request to $75.00/M. Free free API options are available from 11 providers. The page also shows measured API speed and first-token latency.
The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...
Input and output token limits for this model, plus how it ranks on long-context understanding.
Compare Free API pricing across 6 providers. Prices range from $0.0010/request to $75.00/M. CM-API 公益站 offers the lowest rate at $0.0010/request. 11 providers offer free API credits or a free tier.
| Провайдер | Работоспособность | Вариант модели | Группа | Входные данные ($/M) | Выходные данные ($/M) | Скорость (т/с) | Первый токен | Аудит |
|---|---|---|---|---|---|---|---|---|
L1 100% | openrouter/free | default | $0.0050/request | - | — | — | — | |
L1 100% | openrouter/free | free | $0.0010/request | - | — | — | — | |
L1 100% | free | default | $75.00/M | $75.00/M | — | — | — | |
91VIP Бесплатно | L1 0% | openrouter/free | 91vip | Бесплатно | Бесплатно | — | — | — |
L1 0% | openrouter/free | default | $0.010/request | - | — | — | — | |
Astrdark Бесплатно | L1 100% | openrouter/free | default | Бесплатно | Бесплатно | — | — | — |
Futureppo Бесплатно | L1 0% | openrouter/free | 91vip | Бесплатно | Бесплатно | — | — | — |
L1 0% | kilo-auto/free | 2api | $1.00/request | - | — | — | — |
gpt-oss
GPT-OSS is an open-weight language model family designed for self-hosted inference, research, and cost-efficient alternatives to proprietary GPT-class models.
qwen3
Alibaba Qwen3 is the Qwen family's flagship LLM series with dense and MoE variants, seamless thinking/non-thinking modes, and leading open-source performance in math, code, and agent tasks.
glm-4-5-air
Zhipu AI GLM-4.5 Air is a lightweight GLM-4.5 variant optimized for low-latency Chinese and English dialogue, retrieval-augmented apps, and edge deployments.
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
minimax-m2-5
MiniMax M2.5 is MiniMax's flagship text model for coding and agents, with SOTA-level programming and agentic performance, improved token efficiency, and fast high-TPS API deployment.
qwen3-5
Alibaba Qwen3.5 is a Qwen3 generation model with improved reasoning, multilingual support, and efficient inference for chat, coding, and agent applications.