Top for agents
LLM Models
Text-to-text and general language models, separated from media generation categories.
Top for reasoning
Top for coding
Highest throughput
| Model | Context | Input | Output | Providers | Agents | Coding | Reasoning | Knowledge | Math | Multilingual | Multimodal | Instruction following | Throughput | Latency | Release date | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GLM 5.3Z.ai | Context1.0M | Input$1.40/M | Output$4.40/M | Providers +4 | — | 63.3±16.0P | 61.2±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-18 |
|
| Qwen3.8 27BQwen | Context1M | Input$0.425/M | Output$3.10/M | Providers— | 57.2±11.9E | 56.2±10.6E | 58.1±13.9P | 41.5±16.1P | — | — | 57.7±11.3E | 56.1±16.0P | Throughput — | Latency — | Release date2026-08-14 |
|
| Gemini 3.7 FlashGoogle | Context1.0M | Input$0.750/M | Output$3.75/M | Providers +7 | 57.2±8.6 | 58.3±11.9E | 59.3±10.8E | 57.3±14.0P | — | — | 57.2±16.1P | — | Throughput — | Latency — | Release date2026-08-13 |
|
| Qwen3.8 2.4T A95BQwen | Context1.0M | Input$2.00/M | Output$6.00/M | Providers— | — | 60±16.0P | 62.5±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-12 |
|
| Grok 4.6SpaceXAI | Context500K | Input$2.00/M | Output$6.00/M | Providers +15 | 62.8±8.6 | 59.3±11.9E | 59.7±10.8E | 58.4±14.0P | — | — | — | — | Throughput 109 t/s | Latency 10.11s | Release date2026-08-12 |
|
| DeepSeek V4 Pro 0813DeepSeek | Context1.0M | Input— | Output— | Providers— | 57.5±7.7 | 57.1±6.1 | 61±8.3 | 58.4±18.0P | 58.9±12.2E | — | — | 54.9±16.0P | Throughput — | Latency — | Release date2026-08-12 |
|
| Nemotron 3.5 LightningNVIDIA | Context262.1K | Input$0.070/M | Output$0.220/M | Providers +3 | — | 46.4±16.0P | 49±13.9P | — | — | — | — | — | Throughput 213 t/s | Latency 21.68s | Release date2026-08-11 |
|
| Solar Pro 4Upstage | Context524.3K | Input$0.300/M | Output$1.20/M | Providers | — | 54.8±16.0P | 58±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-10 |
|
| Muse Glimmer 30BMeta | Context131.1K | Input— | Output— | Providers— | 47.7±10.2E | 46.6±6.9 | 56.1±10.8E | 43.3±16.0P | 46.1±16.3P | — | 46.4±9.4E | 55.1±16.0P | Throughput — | Latency — | Release date2026-08-09 |
|
| Ling 3.0 TinyinclusionAI | Context262.1K | Input— | Output— | Providers | — | 40.4±16.0P | 48.3±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-06 |
|
| Muse Spark 1.2Meta | Context1.0M | Input$1.25/M | Output$4.25/M | Providers +2 | 56.9±9.3E | 57.5±11.9E | 61.2±10.8E | 55.4±14.0P | — | — | — | — | Throughput 69 t/s | Latency 22.53s | Release date2026-08-05 |
|
| Qwen3.8 MaxQwen | Context1M | Input$2.00/M | Output$6.00/M | Providers +22 | 66.4±10.9E | 66.1±11.2E | 64.4±10.1E | 52.7±14.0P | — | — | 64.7±6.3 | 57.7±16.0P | Throughput 87 t/s | Latency 3.46s | Release date2026-08-03 |
|
| Inkling SmallThinking Machines | Context1.0M | Input$0.300/M | Output$1.20/M | Providers | 55.7±10.9E | 52.7±6.9 | 50.9±8.7 | 52.3±14.0P | 52.5±14.4P | — | 45.9±11.9E | 57.4±16.0P | Throughput — | Latency — | Release date2026-07-30 |
|
| DeepSeek V4 Flash 0731DeepSeek | Context1.3M | Input— | Output— | Providers | 52.1±8.4 | 52.9±6.3 | 59.4±8.3 | 47.5±18.0P | 57.9±12.2E | — | — | — | Throughput 74 t/s | Latency 2.70s | Release date2026-07-31 |
|
| Qwen3.7 FlashQwen | Context1M | Input— | Output— | Providers +4 | — | — | — | — | — | — | — | — | Throughput 388 t/s | Latency 11.58s | Release date2026-07-27 |
|
| Claude Opus 5Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +83 | 66.1±8.2 | 67.7±6.6 | 60.3±8.7 | 67.2±14.0P | — | 64.3±17.1P | 57.9±17.3P | — | Throughput 320 t/s | Latency 4.96s | Release date2026-07-24 |
|
| Ling-3.0-flashInclusionai | Context262.1K | Input$0.075/M | Output$0.220/M | Providers +10 | 47.8±10.2E | 50.1±9.3E | 52.7±10.8E | 38.2±14.0P | 46.7±11.5E | — | — | 54.1±16.0P | Throughput — | Latency — | Release date2026-07-23 |
|
| Laguna S 2.1Poolside | Context1.0M | Input— | Output— | Providers +12 | 53.5±16.0P | 54.1±9.3E | — | — | — | — | — | — | Throughput 41 t/s | Latency 1.46s | Release date2026-07-21 |
|
| Gemini 3.5 Flash-LiteGoogle | Context1.0M | Input$0.300/M | Output$2.50/M | Providers +38 | 43.6±8.8 | 47.2±9.3E | 53±10.8E | 51.4±14.0P | — | — | 45.8±17.3P | — | Throughput — | Latency — | Release date2026-07-21 |
|
| Gemini 3.6 FlashGoogle | Context1.0M | Input$0.750/M | Output$3.75/M | Providers +65 | 53.3±8.6 | 55.6±11.9E | 59.7±10.8E | 57.3±14.0P | — | — | 50.6±17.3P | — | Throughput 737 t/s | Latency 2.34s | Release date2026-07-21 |
|
| InklingThinking Machines | Context1.0M | Input$1.00/M | Output$4.05/M | Providers +12 | 50.9±8.3 | 49.8±6.9 | 55.4±10.8E | 53.1±14.0P | 65.4±16.3P | — | 45.8±11.9E | 56.3±16.0P | Throughput 25 t/s | Latency 0.79s | Release date2026-07-17 |
|
| Muse Spark 1.1Meta | Context1.0M | Input$1.25/M | Output$4.25/M | Providers +2 | 63.1±7.4 | 59±9.3E | 59.5±10.2E | 65.7±14.0P | — | — | 56.8±16.1P | — | Throughput — | Latency — | Release date2026-07-16 |
|
| Kimi K3MoonshotAI | Context1.0M | Input$3.00/M | Output$15.00/M | Providers +82 | 68.3±7.3 | 57.3±11.9E | 59.3±10.8E | 61.4±14.0P | — | — | 65±11.3E | — | Throughput 200 t/s | Latency 33.73s | Release date2026-07-16 |
|
| KAT-Coder-Pro V2.5Kwaipilot | Context256K | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-10 |
|
| GPT-5.6 Sol ProOpenAI | Context1.1M | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-09 |
|
| GPT-5.6 Terra ProOpenAI | Context1.1M | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-09 |
|
| GPT-5.6 Luna ProOpenAI | Context1.1M | Input— | Output— | Providers +3 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-09 |
|
| GPT-3.5 Turbo 16kOpenAI | Context16.4K | Input— | Output— | Providers +1 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2023-08-28 |
|
| GPT-4 Turbo PreviewOpenAI | Context128K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-01-25 |
|
| GPT-3.5 Turbo (older v0613)OpenAI | Context4.1K | Input— | Output— | Providers +1 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-01-25 |
|
| Llama 3 8B InstructMeta | Context8.2K | Input— | Output— | Providers +1 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| GPT-4o (2024-05-13)OpenAI | Context128K | Input$5.00/M | Output$15.00/M | Providers | — | 43.5±13.9P | 42.7±11.1E | — | 46±11.8E | — | — | — | Throughput — | Latency — | Release date2024-05-13 |
|
| GPT-4o-mini (2024-07-18)OpenAI | Context128K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-07-18 |
|
| Llama 3.1 70B InstructMeta | Context131.1K | Input— | Output— | Providers +1 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-07-23 |
|
| Llama 3.1 8B InstructMeta | Context131.1K | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-07-23 |
|
| GPT-4o (2024-08-06)OpenAI | Context128K | Input$2.50/M | Output$10.00/M | Providers | — | 44.3±13.9P | 39.2±13.9P | — | 46.2±11.8E | — | — | — | Throughput — | Latency — | Release date2024-08-06 |
|
| Hermes 3 70B InstructNous | Context131.1K | Input$0.700/M | Output$0.700/M | Providers— | — | 36.6±13.9P | 37.8±11.1E | — | 39.6±11.8E | — | — | — | Throughput — | Latency — | Release date2024-08-18 |
|
| Llama 3.2 11B Vision InstructMeta | Context131.1K | Input— | Output— | Providers +4 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-09-25 |
|
| Llama 3.2 1B InstructMeta | Context60K | Input— | Output— | Providers +4 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-09-25 |
|
| Llama 3.2 3B InstructMeta | Context131.1K | Input— | Output— | Providers +8 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-09-25 |
|
| Mistral Large 2407Mistralai | Context131.1K | Input$2.00/M | Output$6.00/M | Providers | — | 40.4±13.9P | 41.1±11.1E | — | 44.5±11.8E | — | — | — | Throughput — | Latency — | Release date2024-11-19 |
|
| GPT-4o (2024-11-20)OpenAI | Context128K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-11-20 |
|
| Llama 3.3 70B InstructMeta | Context131.1K | Input— | Output— | Providers +6 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-12-06 |
|
| DeepSeek V3DeepSeek | Context163.8K | Input— | Output— | Providers +15 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-12-26 |
|
| R1 Distill Llama 70BDeepSeek | Context8.2K | Input$0.700/M | Output$1.10/M | Providers | — | 42.6±13.9P | 45.3±11.1E | — | 55±11.8E | — | — | — | Throughput — | Latency — | Release date2025-01-23 |
|
| Qwen2.5 VL 72B InstructQwen | Context128K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-02-01 |
|
| o3 Mini HighOpenAI | Context200K | Input$1.10/M | Output$4.40/M | Providers | — | 53.3±13.9P | 51±11.1E | — | 61.3±11.8E | — | — | — | Throughput — | Latency — | Release date2025-02-12 |
|
| Gemma 3 27BGoogle | Context262.1K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-03-12 |
|
| GPT-4o Search PreviewOpenAI | Context128K | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-03-12 |
|
| GPT-4o-mini Search PreviewOpenAI | Context128K | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-03-12 |
|
| Gemma 3 12BGoogle | Context131.1K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-03-13 |
|
| Gemma 3 4BGoogle | Context131.1K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-03-13 |
|
| Mistral Small 3.1 24BMistral | Context128K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-03-17 |
|
| o4 Mini HighOpenAI | Context200K | Input— | Output— | Providers— | — | — | — | — | 51.7±16.1P | — | — | — | Throughput — | Latency — | Release date2025-04-16 |
|
| Qwen3 235B A22BQwen | Context131.1K | Input$0.700/M | Output$8.40/M | Providers | — | 43.1±13.9P | 45.7±11.1E | — | 51.1±11.8E | — | — | — | Throughput — | Latency — | Release date2025-04-28 |
|
| Qwen3 32BQwen | Context131.1K | Input$0.160/M | Output$0.640/M | Providers | — | 41.3±13.9P | 46.4±11.1E | — | 49.9±11.8E | — | — | — | Throughput — | Latency — | Release date2025-04-28 |
|
| Qwen3 14BQwen | Context131.1K | Input$0.350/M | Output$4.20/M | Providers | — | 40.3±13.9P | 44.8±11.1E | — | 49.8±11.8E | — | — | — | Throughput — | Latency — | Release date2025-04-28 |
|
| Qwen3 8BQwen | Context131.1K | Input$0.180/M | Output$2.10/M | Providers | — | 32.6±13.9P | 42.2±11.1E | — | 48.5±11.8E | — | — | — | Throughput — | Latency — | Release date2025-04-28 |
|
| Qwen3 30B A3BQwen | Context131.1K | Input$0.200/M | Output$2.40/M | Providers | — | 40.9±13.9P | 43.2±11.1E | — | 49.4±11.8E | — | — | — | Throughput — | Latency — | Release date2025-04-28 |
|
| Llama Guard 4 12BMeta | Context1.0M | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-04-30 |
|
