Choose a model to open its comparison page.
The page also shows measured API speed and first-token latency.
DeepSeek Chat V3.1 is an instruction-tuned variant in the DeepSeek series, optimized for following instructions and conversational tasks.
Input and output token limits for this model, plus how it ranks on long-context understanding.
Сторонние данные эндпоинтов OpenRouter показаны отдельно от измерений LMSpeed. Некоторые 30-минутные live-метрики появляются только после синхронизации с OpenRouter API key.
| Provider endpoint | Вход | Вывод | Доступность 1 день | Задержка 30м | Пропускная способность 30м | Контекст / вывод |
|---|---|---|---|---|---|---|
Novita novita/fp8 | $0.270/M | $1/M | 100.0% | — | — | 131.1K токенов / 32.8K токенов |
AtlasCloud atlas-cloud/fp8 | $0.300/M | $0.950/M | 99.8% | — | — | 131.1K токенов / 65.5K токенов |
SambaNova sambanova/fp8 | $0.650/M | $1.50/M | 99.8% | — | — | 131.1K токенов / 7.2K токенов |
DeepInfra deepinfra/fp4 | $0.210/M | $0.790/M | 99.7% | — | — | 163.8K токенов / 32.8K токенов |
Google google-vertex/us-west2 | $0.600/M | $1.70/M | 99.6% | — | — | 163.8K токенов / 32.8K токенов |
WandB wandb/fp8 | $0.550/M | $1.65/M | 98.6% | — | — | 161K токенов / 161K токенов |
SiliconFlow siliconflow/fp8 | $0.270/M | $1/M | 97.2% | — | — | 163.8K токенов / 163.8K токенов |
Mara mara | $0.600/M | $1.70/M | 88.1% | — | — | 131.1K токенов / 7.2K токенов |
gemini-2-5-pro
Google Gemini 2.5 Pro is Google advanced multimodal model with a 1M-token context window, strong STEM reasoning, and native support for images, audio, and video understanding.
qwen3
Alibaba Qwen3 is the Qwen family's flagship LLM series with dense and MoE variants, seamless thinking/non-thinking modes, and leading open-source performance in math, code, and agent tasks.
kimi-k2
Moonshot Kimi K2 is a trillion-parameter MoE model from Moonshot AI, built for long-context chat, agentic tool use, and high-quality Chinese-English bilingual generation.
qwen3-coder
Alibaba Qwen3 Coder is a code-specialized variant in the Qwen series, optimized for code generation, debugging, and software development tasks.
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
gpt-5-3-codex
OpenAI GPT-5.3 Codex is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks.