Top for agents
LLM Models
Text-to-text and general language models, separated from media generation categories.
Top for reasoning
Top for coding
Highest throughput
| Model | Context | Input | Output | Providers | Agents | Coding | Reasoning | Knowledge | Math | Multilingual | Multimodal | Instruction following | Throughput | Latency | Release date | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V3 | Context— | Input$0.360/M | Output$0.890/M | Providers +150 | 34.6±11.8E | 37.4±11.2E | 44.4±8.7 | 41.1±16.0P | 49.5±9.3E | — | — | 39.7±11.9E | Throughput 49 t/s | Latency 3.39s | Release date— |
|
| Qwen3 | Context262.1K | Input$0.200/M | Output$0.800/M | Providers +126 | — | 45.7±13.9P | 47.6±11.1E | 36.9±16.3P | 58.3±11.8E | 36.7±16.4P | — | — | Throughput 100 t/s | Latency 12.55s | Release date2025-07-21 |
|
| Qwen3 VL | Context— | Input— | Output— | Providers +17 | — | — | — | — | — | — | — | — | Throughput 206 t/s | Latency 6.63s | Release date— |
|
| Devstral 2 | Context— | Input— | Output— | Providers +4 | — | 46.2±13.9P | 45.3±11.1E | — | 43.2±16.1P | — | — | — | Throughput 71 t/s | Latency 1.34s | Release date— |
|
| Ministral 3 | Context— | Input$0.200/M | Output$0.200/M | Providers +2 | — | 31.5±13.9P | 36.6±11.1E | — | 40.5±16.1P | — | — | — | Throughput 208 t/s | Latency 2.84s | Release date— |
|
| Mistral Large 3 | Context— | Input$0.500/M | Output$1.50/M | Providers +7 | 39.3±11.8E | 47.9±13.9P | 44.5±8.7 | 41.6±16.0P | 43.5±16.1P | — | — | 42.2±16.0P | Throughput 29 t/s | Latency 3.00s | Release date— |
|
| Qwen3 CoderQwen | Context262.1K | Input— | Output— | Providers +61 | — | — | — | — | — | — | — | — | Throughput 243 t/s | Latency 5.28s | Release date2025-07-23 |
|
| Qwen3 Embedding | Context— | Input— | Output— | Providers +42 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| Qwen3 Next | Context— | Input— | Output— | Providers +7 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| DeepSeek Coder | Context— | Input— | Output— | Providers +8 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| Qwen Image | Context— | Input— | Output— | Providers +34 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| Phi 4Microsoft | Context16.4K | Input$0.125/M | Output$0.500/M | Providers +17 | 28.4±16.0P | 39.2±13.9P | 39.1±8.7 | 37.5±16.0P | 46.9±11.8E | — | — | 37.8±16.0P | Throughput 43 t/s | Latency 1.32s | Release date2025-01-10 |
|
| Intern-S1 | Context— | Input— | Output— | Providers +4 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| MiniMax M1MiniMax | Context1M | Input$0.550/M | Output$2.20/M | Providers +30 | — | 51.2±13.9P | 49.4±11.1E | — | 58.9±11.8E | — | — | — | Throughput — | Latency — | Release date2025-06-17 |
|
| MAI-DS-R1 | Context— | Input— | Output— | Providers +10 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| DeepSeek VL2 | Context— | Input— | Output— | Providers +8 | — | — | — | — | — | — | — | — | Throughput 126 t/s | Latency 0.75s | Release date— |
|
| GLM-4 | Context— | Input— | Output— | Providers +56 | — | — | — | — | — | — | — | — | Throughput 61 t/s | Latency 0.87s | Release date— |
|
| GPT-OSS | Context131.1K | Input— | Output— | Providers +137 | — | — | — | — | — | — | — | — | Throughput 383 t/s | Latency 2.36s | Release date2025-08-05 |
|
| GPT-4oOpenAI | Context128K | Input$2.50/M | Output$10.00/M | Providers +132 | 39±16.0P | 46±13.9P | 39.9±8.7 | 49±16.0P | 42.2±9.3E | — | — | 41.6±16.0P | Throughput 82 t/s | Latency 3.44s | Release date2024-11-20 |
|
| MiniMax M2.5MiniMax | Context204.8K | Input$0.300/M | Output$1.20/M | Providers +186 | — | 50.2±11.8E | 54.6±13.9P | — | — | — | — | — | Throughput 61 t/s | Latency 9.68s | Release date2026-02-12 |
|
| GPT-5.4 MiniOpenAI | Context400K | Input$0.750/M | Output$4.50/M | Providers +238 | 51±8.1 | 57.5±11.8E | 55.6±10.8E | 46.6±14.0P | 49.7±16.1P | — | 47.1±16.1P | 53.6±16.0P | Throughput 134 t/s | Latency 3.54s | Release date2026-03-17 |
|
| MiniMax Hailuo 2.3 | Context— | Input— | Output— | Providers +16 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| GLM-5Z.ai | Context204.8K | Input$1.00/M | Output$3.20/M | Providers +185 | 41.1±5.6 | 51.5±8.1 | 48.9±8.1 | 43.3±16.3P | 51.1±9.1E | 43.8±12.2E | — | 52.4±11.9E | Throughput 51 t/s | Latency 21.59s | Release date2026-02-11 |
|
| GPT-4o MiniOpenAI | Context128K | Input$0.150/M | Output$0.600/M | Providers +121 | 30.8±16.0P | 37.4±13.9P | 40.3±11.1E | 53.4±16.0P | 46.1±11.8E | — | — | 40.5±16.0P | Throughput 84 t/s | Latency 4.10s | Release date2024-07-18 |
|
| Claude Sonnet 4.5Anthropic | Context1M | Input$3.00/M | Output$15.00/M | Providers +174 | 46.6±8.7 | 51.8±11.2E | 45.6±8.9E | 54.1±16.1P | 42.5±12.3E | — | — | — | Throughput 41 t/s | Latency 4.57s | Release date2025-09-29 |
|
| GPT-5.1OpenAI | Context400K | Input$1.25/M | Output$10.00/M | Providers +153 | 47.7±9.3E | 49.1±11.2E | 54±8.7 | 53±16.0P | 54±16.1P | — | — | 53.5±16.0P | Throughput 142 t/s | Latency 2.78s | Release date2025-11-13 |
|
| GLM-4.6VZ.ai | Context131.1K | Input$0.300/M | Output$0.900/M | Providers +60 | — | 40.1±13.9P | 49.8±11.1E | — | 52.2±16.1P | — | — | — | Throughput 40 t/s | Latency 27.37s | Release date2025-12-08 |
|
| Gemini 2.5 ProGoogle | Context1.0M | Input$1.25/M | Output$10.00/M | Providers +171 | 43.6±9.3E | 39.1±8.9 | 54.2±8.7 | 35±14.0P | 56.1±9.3E | — | — | 45.9±16.0P | Throughput 88 t/s | Latency 16.94s | Release date2025-06-17 |
|
| O1OpenAI | Context200K | Input$15.00/M | Output$60.00/M | Providers +75 | 45.6±16.0P | 50.6±13.9P | 50.7±8.7 | 62.5±17.4P | 52.6±9.3E | — | — | 51.5±11.9E | Throughput — | Latency — | Release date2024-12-17 |
|
| O1 MiniOpenAI | Context— | Input— | Output— | Providers +60 | — | 47.4±13.9P | 44.7±11.1E | — | 54.9±11.8E | — | — | — | Throughput 32 t/s | Latency 13.84s | Release date— |
|
| GPT-4.1OpenAI | Context1.0M | Input$2.00/M | Output$8.00/M | Providers +109 | 40±11.8E | 38.3±11.2E | 48.6±8.7 | 57±17.4P | 49±9.3E | — | — | 42.1±11.9E | Throughput 85 t/s | Latency 1.92s | Release date2025-04-14 |
|
| O3OpenAI | Context200K | Input$2.00/M | Output$8.00/M | Providers +85 | 49.3±16.0P | 55.2±13.9P | 54.3±8.7 | 47.7±16.0P | 58.8±9.3E | — | — | 53±16.0P | Throughput 134 t/s | Latency 2.82s | Release date2025-04-16 |
|
| Kimi K2MoonshotAI | Context131.1K | Input$0.570/M | Output$2.30/M | Providers +87 | 45.3±16.0P | 48.3±13.9P | 48.6±8.7 | 44.5±16.0P | 53.7±9.3E | — | — | 43.8±16.0P | Throughput 29 t/s | Latency 2.21s | Release date2025-09-04 |
|
| Claude Opus 4.6Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +228 | 54.9±5.7 | 58.9±8.0 | 53.4±8.7 | 53.7±13.7P | 56.5±16.1P | 54±17.3P | 48.4±11.1E | 44.7±16.0P | Throughput 45 t/s | Latency 5.30s | Release date2026-02-04 |
|
| Gemini 3 ProGoogle | Context— | Input$2.00/M | Output$12.00/M | Providers +139 | 49±9.3E | 55.5±11.2E | 54.1±6.4 | 56.1±14.0P | 55.7±16.1P | 54±17.3P | 49.9±6.5 | 52.6±16.0P | Throughput 73 t/s | Latency 10.59s | Release date— |
|
| Phi 4 Mini Instruct | Context131.1K | Input— | Output— | Providers +28 | — | 27.9±13.9P | 34.3±11.1E | — | 42±11.8E | — | — | — | Throughput — | Latency — | Release date2025-10-17 |
|
| GPT-5 MiniOpenAI | Context400K | Input$0.250/M | Output$2.00/M | Providers +111 | — | 49.5±11.2E | 53.5±11.1E | — | 52.2±16.1P | — | — | — | Throughput 90 t/s | Latency 7.03s | Release date2025-08-07 |
|
| GPT-5.1 Codex MaxOpenAI | Context400K | Input— | Output— | Providers +126 | 49.9±16.0P | 50.8±11.8E | 54.3±10.8E | 50±16.0P | — | — | — | 52.5±16.0P | Throughput 56 t/s | Latency 6.29s | Release date2025-12-04 |
|
| Gemini 1.5 ProGoogle | Context— | Input— | Output— | Providers +21 | — | 40.2±13.9P | 39.7±11.1E | 54.2±16.0P | 43.7±11.8E | — | — | — | Throughput 18 t/s | Latency 2.58s | Release date— |
|
| GPT-4.1 MiniOpenAI | Context1.0M | Input$0.400/M | Output$1.60/M | Providers +112 | 40.3±11.8E | 39±11.2E | 45.4±8.7 | 49.3±17.4P | 49.9±9.3E | — | — | 42.4±11.9E | Throughput 93 t/s | Latency 2.69s | Release date2025-04-14 |
|
| Grok Imagine VideoSpaceXAI | Context0 | Input— | Output— | Providers +21 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-05-18 |
|
| Claude Opus 4Anthropic | Context200K | Input$15.00/M | Output$75.00/M | Providers +91 | — | 51±13.9P | 51.7±11.1E | — | 54.4±11.8E | — | — | — | Throughput 25 t/s | Latency 12.05s | Release date2025-05-22 |
|
| Gemini 2.5 Flash ImageGoogle | Context32.8K | Input— | Output— | Providers +93 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-10-07 |
|
| GPT-5.4OpenAI | Context1.1M | Input$2.50/M | Output$15.00/M | Providers +292 | 57.2±5.1 | 62.4±10.3E | 56.2±8.7 | 59.9±15.1P | 57.6±16.1P | — | 52.9±6.6 | 53.9±16.0P | Throughput 49 t/s | Latency 4.45s | Release date2026-03-05 |
|
| GPT Audio MiniOpenAI | Context128K | Input— | Output— | Providers +16 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-01-19 |
|
| GLM-4.7Z.ai | Context204.8K | Input$0.600/M | Output$2.20/M | Providers +151 | 46.7±8.6 | 53±10.5E | 54.4±8.7 | 34.7±14.0P | 48.8±12.3E | — | — | 51.8±16.0P | Throughput 94 t/s | Latency 18.83s | Release date2025-12-22 |
|
| Claude Opus 4.5Anthropic | Context200K | Input$5.00/M | Output$25.00/M | Providers +166 | 48.5±5.3 | 55.7±8.1 | 60.6±8.1 | 57.6±11.6E | 49.9±9.1E | 51.7±12.2E | 32.4±6.5 | 45.7±11.9E | Throughput 54 t/s | Latency 3.12s | Release date2025-11-24 |
|
| GLM-4.6Z.ai | Context204.8K | Input$0.550/M | Output$2.20/M | Providers +101 | 48.4±16.0P | 41.9±11.2E | 42.9±8.7 | 43.6±16.0P | 44.8±16.1P | — | — | 42.4±16.0P | Throughput 53 t/s | Latency 16.68s | Release date2025-09-30 |
|
| Gemini 2.5 Flash LiteGoogle | Context1.0M | Input$0.100/M | Output$0.400/M | Providers +128 | — | 36.5±13.9P | 42.9±11.1E | — | 53.3±11.8E | — | — | — | Throughput 277 t/s | Latency 1.41s | Release date2025-09-25 |
|
| DeepSeek R1DeepSeek | Context64K | Input$1.35/M | Output$3.00/M | Providers +149 | 41.2±16.0P | 49.7±13.9P | 50.2±8.7 | 44.7±16.0P | 57±11.8E | — | — | 43.3±16.0P | Throughput 52 t/s | Latency 10.36s | Release date2025-05-28 |
|
| Devstral SmallMistral AI | Context— | Input— | Output— | Providers +7 | — | 38.8±13.9P | 39.5±11.1E | — | 43.4±11.8E | — | — | — | Throughput — | Latency — | Release date— |
|
| GPT-3.5 Turbo InstructOpenAI | Context4.1K | Input— | Output— | Providers +44 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2023-09-28 |
|
| GPT-3.5 TurboOpenAI | Context16.4K | Input$0.500/M | Output$1.50/M | Providers +98 | — | 48.8±18.4P | 33.5±11.8E | — | 37.8±16.0P | — | — | — | Throughput 115 t/s | Latency 2.21s | Release date2024-01-25 |
|
| Gemini 3.1 Flash LiteGoogle | Context1.0M | Input— | Output— | Providers +128 | 44.5±16.1P | 27.5±16.1P | — | — | — | — | 42.6±16.1P | — | Throughput 191 t/s | Latency 6.58s | Release date2026-05-07 |
|
| O3 ProOpenAI | Context200K | Input$20.00/M | Output$80.00/M | Providers +61 | — | — | 54±16.0P | 60.1±16.0P | — | — | — | — | Throughput — | Latency — | Release date2025-06-10 |
|
| GPT-4 TurboOpenAI | Context128K | Input$10.00/M | Output$30.00/M | Providers +73 | — | 43.3±13.9P | 43.2±11.8E | 53.6±16.0P | 45.8±11.8E | — | — | — | Throughput 50 t/s | Latency 0.97s | Release date2024-04-09 |
|
| GPT-5.1 Codex MiniOpenAI | Context400K | Input$0.250/M | Output$2.00/M | Providers +125 | — | 56.6±13.9P | 53.2±11.1E | — | 53.4±16.1P | — | — | — | Throughput 136 t/s | Latency 4.28s | Release date2025-11-13 |
|
| Claude Opus 4.1Anthropic | Context200K | Input$15.00/M | Output$75.00/M | Providers +100 | 46.5±16.2P | 47.1±16.1P | — | 59±16.0P | — | — | — | — | Throughput 22 t/s | Latency 3.56s | Release date2025-08-05 |
|
| Qwen3 Omni Flash | Context— | Input— | Output— | Providers +38 | — | — | — | — | — | — | — | — | Throughput 62 t/s | Latency 5.13s | Release date— |
|
| DeepSeek ReasonerDeepSeek | Context— | Input— | Output— | Providers +85 | — | — | — | — | — | — | — | — | Throughput 28 t/s | Latency 6.40s | Release date— |
|
