Top for agents
LLM Models
Text-to-text and general language models, separated from media generation categories.
Top for reasoning
Top for coding
Highest throughput
| Model | Context | Input | Output | Providers | Agents | Coding | Reasoning | Knowledge | Math | Multilingual | Multimodal | Instruction following | Throughput | Latency | Release date | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Ling 3.0 Flash VLinclusionAI | Context131.1K | Input— | Output— | Providers | — | 51.9±16.0P | 55.1±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-09-10 |
|
| DeepSeek V4.1 FlashDeepSeek | Context1.0M | Input$0.300/M | Output$1.20/M | Providers +42 | 50.6±18.4P | 57.6±16.0P | 56.9±16.0P | — | — | — | — | — | Throughput — | Latency — | Release date2026-09-10 |
|
| Nex-N2.5-ProNex AGI | Context262.1K | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-09-08 |
|
| Nex-N2.5-MiniNex AGI | Context262.1K | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-09-08 |
|
| Qwen3.8 Max (0902)Qwen | Context1M | Input— | Output— | Providers +1 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-09-03 |
|
| GPT-6 AstraOpenAI | Context1.1M | Input$10.00/M | Output$50.00/M | Providers +88 | 63±8.4 | 59.5±11.9E | 62.1±8.7 | 58±16.6P | 72.3±16.1P | — | 66.4±16.4P | — | Throughput — | Latency — | Release date2026-09-04 |
|
| Ling 3.0 Flash SanteinclusionAI | Context262.1K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-09-04 |
|
| Muse Spark 1.3 ContributorMeta | Context1.0M | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-09-02 |
|
| Muse Spark 1.3Meta | Context1.0M | Input$1.25/M | Output$4.25/M | Providers +2 | 63.4±8.4 | 58.3±11.9E | 60.9±10.8E | 56.9±14.0P | — | — | — | — | Throughput — | Latency — | Release date2026-09-02 |
|
| Gemini 3.8 FlashGoogle | Context1.0M | Input$0.750/M | Output$3.75/M | Providers +60 | 59.9±11.3E | 58.8±11.9E | 59.7±10.8E | 61.3±14.0P | — | — | 54±16.4P | — | Throughput — | Latency — | Release date2026-09-02 |
|
| Claude Fable 5.1Anthropic | Context1M | Input$10.00/M | Output$50.00/M | Providers +66 | 61.7±8.6 | 68.5±6.9 | 61.2±8.7 | 68±14.0P | — | — | — | — | Throughput — | Latency — | Release date2026-09-01 |
|
| Granite 4.2 8BIBM | Context131.1K | Input$0.060/M | Output$0.250/M | Providers— | — | 42±16.0P | 45.9±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-31 |
|
| Ling 3.0 Flash FininclusionAI | Context262.1K | Input— | Output— | Providers +6 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-27 |
|
| Qwen3.8 FlashQwen | Context1M | Input— | Output— | Providers +26 | — | — | — | — | — | — | — | — | Throughput 126 t/s | Latency 2.96s | Release date2026-08-26 |
|
| GLM 5.3 FlashZ.ai | Context1.3M | Input$0.150/M | Output$0.500/M | Providers +79 | 60.6±11.9E | 56.6±9.3E | 57.6±11.1E | 54.2±14.0P | — | — | 57.9±16.1P | — | Throughput 336 t/s | Latency 31.71s | Release date2026-08-26 |
|
| Muse Spark 1.2 ContributorMeta | Context1.0M | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-21 |
|
| DeepSeek V4 Flash Vision ExpDeepSeek | Context1.0M | Input— | Output— | Providers +24 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-21 |
|
| Ox AlphaStealth | Context1.0M | Input— | Output— | Providers +1 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-20 |
|
| GLM 5.3Z.ai | Context1.3M | Input$1.40/M | Output$4.40/M | Providers +80 | 50.8±18.4P | 62.9±16.0P | 60.8±13.9P | — | — | — | — | — | Throughput 379 t/s | Latency 30.00s | Release date2026-08-18 |
|
| Qwen3.8 27BQwen | Context1M | Input$0.500/M | Output$3.00/M | Providers +16 | 58.7±8.4 | 51.6±8.6 | 54.7±10.8E | 43.9±14.0P | — | — | 57.6±11.3E | 56.3±16.0P | Throughput — | Latency — | Release date2026-08-14 |
|
| Gemini 3.7 FlashGoogle | Context1.0M | Input$0.750/M | Output$3.75/M | Providers +70 | 54.6±8.6 | 57.9±11.9E | 58.5±10.8E | 60.8±14.0P | — | — | 56.9±16.1P | — | Throughput — | Latency — | Release date2026-08-13 |
|
| Qwen3.8 2.4T A95BQwen | Context1.0M | Input$2.00/M | Output$6.00/M | Providers | — | 59.2±16.0P | 62.1±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-12 |
|
| Seed 2.1 TurboByteDance Seed | Context262.1K | Input— | Output— | Providers +1 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-10 |
|
| Grok 4.6SpaceXAI | Context500K | Input$2.00/M | Output$6.00/M | Providers +92 | 60.8±8.4 | 58.5±11.9E | 59.5±10.8E | 60.2±14.0P | — | — | — | — | Throughput 150 t/s | Latency 22.76s | Release date2026-08-12 |
|
| DeepSeek V4 Pro 0813DeepSeek | Context1.0M | Input— | Output— | Providers +29 | 57.7±7.3 | 54.9±6.1 | 60.8±8.3 | 58.4±18.0P | 58.9±12.2E | — | — | 55±16.0P | Throughput — | Latency — | Release date2026-08-12 |
|
| LFM2.5-2.6BLiquidAI | Context65.5K | Input— | Output— | Providers +4 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-11 |
|
| Nemotron 3.5 LightningNVIDIA | Context262.1K | Input$0.060/M | Output$0.200/M | Providers +16 | — | 42.5±16.0P | 48.7±13.9P | — | — | — | — | — | Throughput 391 t/s | Latency 34.71s | Release date2026-08-11 |
|
| Solar Pro 4Upstage | Context524.3K | Input$0.300/M | Output$1.20/M | Providers +1 | — | 52.2±16.0P | 57.6±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-10 |
|
| Muse Glimmer 30BMeta | Context131.1K | Input— | Output— | Providers +5 | 46.2±9.9E | 43.6±6.9 | 55.9±10.8E | 41.7±14.0P | 46.1±16.3P | — | 45.9±9.4E | 55.2±16.0P | Throughput — | Latency — | Release date2026-08-09 |
|
| Ling 3.0 TinyinclusionAI | Context262.1K | Input— | Output— | Providers +1 | — | 35.3±16.0P | 48±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-06 |
|
| Muse Spark 1.2Meta | Context1.0M | Input$1.25/M | Output$4.25/M | Providers +2 | 56.3±8.9 | 55.7±11.9E | 58.5±10.8E | 58.8±14.0P | — | — | — | — | Throughput 69 t/s | Latency 22.53s | Release date2026-08-05 |
|
| Qwen3.8 MaxQwen | Context1M | Input$2.00/M | Output$6.00/M | Providers +65 | 61.5±10.9E | 61.7±11.2E | 63.5±10.1E | 52.2±14.0P | — | — | 64.3±6.3 | 57.9±16.0P | Throughput 292 t/s | Latency 4.31s | Release date2026-08-03 |
|
| Inkling SmallThinking Machines | Context1.0M | Input$0.300/M | Output$1.20/M | Providers +7 | 54.6±10.8E | 51.9±6.9 | 50.9±8.7 | 49.7±14.0P | 52.5±14.4P | — | 45.8±11.9E | 57.6±16.0P | Throughput — | Latency — | Release date2026-07-30 |
|
| DeepSeek V4 Flash 0731DeepSeek | Context1.3M | Input— | Output— | Providers +34 | 51.3±8.3 | 50.6±6.3 | 58.7±8.3 | 47.5±18.0P | 57.9±12.2E | — | — | — | Throughput 74 t/s | Latency 2.70s | Release date2026-07-31 |
|
| Qwen3.7 FlashQwen | Context1M | Input— | Output— | Providers +18 | — | — | — | — | — | — | — | — | Throughput 388 t/s | Latency 11.58s | Release date2026-07-27 |
|
| Claude Opus 5Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +107 | 65.9±7.7 | 67.5±6.6 | 59.5±8.7 | 66.9±14.0P | — | 64.3±17.1P | 56.7±16.4P | — | Throughput 284 t/s | Latency 5.44s | Release date2026-07-24 |
|
| Ling-3.0-flashinclusionAI | Context262.1K | Input$0.075/M | Output$0.220/M | Providers +11 | 46.7±10.2E | 48.4±9.3E | 53.1±10.8E | 38.9±14.0P | 46.7±11.5E | — | — | 54.2±16.0P | Throughput — | Latency — | Release date2026-07-23 |
|
| Laguna S 2.1Poolside | Context1.0M | Input— | Output— | Providers +18 | 52.7±16.0P | 53.2±9.3E | — | — | — | — | — | — | Throughput 41 t/s | Latency 1.46s | Release date2026-07-21 |
|
| Gemini 3.5 Flash-LiteGoogle | Context1.0M | Input$0.300/M | Output$2.50/M | Providers +49 | 42.1±8.7 | 44.5±9.3E | 52±10.8E | 51±14.0P | — | — | 44.7±16.4P | — | Throughput — | Latency — | Release date2026-07-21 |
|
| Gemini 3.6 FlashGoogle | Context1.0M | Input$0.750/M | Output$3.75/M | Providers +87 | 51.4±8.6 | 53.1±11.9E | 58.6±10.8E | 59±14.0P | — | — | 53.9±16.4P | — | Throughput 737 t/s | Latency 2.34s | Release date2026-07-21 |
|
| InklingThinking Machines | Context1.0M | Input$1.00/M | Output$4.05/M | Providers +19 | 48.6±8.2 | 47.9±6.9 | 55.2±10.8E | 51.3±14.0P | 65.4±16.3P | — | 45.7±11.9E | 56.4±16.0P | Throughput 25 t/s | Latency 0.79s | Release date2026-07-17 |
|
| Muse Spark 1.1Meta | Context1.0M | Input$1.25/M | Output$4.25/M | Providers +2 | 61.6±7.4 | 58±9.3E | 57.5±10.2E | 63.3±14.0P | — | — | 56.5±16.1P | — | Throughput — | Latency — | Release date2026-07-16 |
|
| Kimi K3MoonshotAI | Context1.0M | Input$3.00/M | Output$15.00/M | Providers +115 | 65.7±7.3 | 56.4±11.9E | 60.1±10.8E | 59.4±14.0P | — | — | 64.9±11.3E | — | Throughput 138 t/s | Latency 30.06s | Release date2026-07-16 |
|
| GPT-5.6 Sol ProOpenAI | Context1.1M | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-09 |
|
| GPT-5.6 Terra ProOpenAI | Context1.1M | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-09 |
|
| GPT-5.6 Luna ProOpenAI | Context1.1M | Input— | Output— | Providers +4 | — | — | — | — | — | — | — | — | Throughput 193 t/s | Latency 15.30s | Release date2026-07-09 |
|
| GPT-3.5 Turbo 16kOpenAI | Context16.4K | Input— | Output— | Providers +3 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2023-08-28 |
|
| GPT-4 Turbo PreviewOpenAI | Context128K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-01-25 |
|
| GPT-3.5 Turbo (older v0613)OpenAI | Context4.1K | Input— | Output— | Providers +3 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-01-25 |
|
| Llama 3 8B InstructMeta | Context8.2K | Input— | Output— | Providers +1 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| GPT-4o (2024-05-13)OpenAI | Context128K | Input$5.00/M | Output$15.00/M | Providers +2 | — | 44.7±16.0P | 42.6±11.1E | — | 46±11.8E | — | — | — | Throughput — | Latency — | Release date2024-05-13 |
|
| GPT-4o-mini (2024-07-18)OpenAI | Context128K | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-07-18 |
|
| Llama 3.1 70B InstructMeta | Context131.1K | Input— | Output— | Providers +1 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-07-23 |
|
| Llama 3.1 8B InstructMeta | Context131.1K | Input— | Output— | Providers +3 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-07-23 |
|
| GPT-4o (2024-08-06)OpenAI | Context128K | Input$2.50/M | Output$10.00/M | Providers +1 | — | 44.3±16.0P | 38.9±13.9P | — | 46.2±11.8E | — | — | — | Throughput — | Latency — | Release date2024-08-06 |
|
| Hermes 3 70B InstructNous | Context131.1K | Input$0.700/M | Output$0.700/M | Providers— | — | 40.9±16.0P | 37.6±11.1E | — | 39.6±11.8E | — | — | — | Throughput — | Latency — | Release date2024-08-18 |
|
| Llama 3.2 11B Vision InstructMeta | Context131.1K | Input— | Output— | Providers +4 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-09-25 |
|
| Llama 3.2 1B InstructMeta | Context60K | Input— | Output— | Providers +4 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-09-25 |
|
| Llama 3.2 3B InstructMeta | Context131.1K | Input— | Output— | Providers +8 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-09-25 |
|
| Mistral Large 2407Mistralai | Context131.1K | Input$2.00/M | Output$6.00/M | Providers | — | 43.1±16.0P | 41±11.1E | — | 44.5±11.8E | — | — | — | Throughput — | Latency — | Release date2024-11-19 |
|
