ModelContextInputOutputProvidersAgentsCodingReasoningKnowledgeMathMultilingualMultimodalInstruction followingThroughputLatencyRelease date
GLM 5.3Z.aiContext1.0MInput$1.40/MOutput$4.40/MProviders
+4
63.3±16.0P
61.2±13.9P
Throughput
Latency
Release date2026-08-18
  • 63.3
  • 61.2
Qwen3.8 27BQwenContext1MInput$0.425/MOutput$3.10/MProviders
57.2±11.9E
56.2±10.6E
58.1±13.9P
41.5±16.1P
57.7±11.3E
56.1±16.0P
Throughput
Latency
Release date2026-08-14
  • 57.2
  • 56.2
  • 58.1
  • 41.5
  • 57.7
  • 56.1
Gemini 3.7 FlashGoogleContext1.0MInput$0.750/MOutput$3.75/MProviders
+7
57.2±8.6
58.3±11.9E
59.3±10.8E
57.3±14.0P
57.2±16.1P
Throughput
Latency
Release date2026-08-13
  • 57.2
  • 58.3
  • 59.3
  • 57.3
  • 57.2
Qwen3.8 2.4T A95BQwenContext1.0MInput$2.00/MOutput$6.00/MProviders
60±16.0P
62.5±13.9P
Throughput
Latency
Release date2026-08-12
  • 60
  • 62.5
Grok 4.6SpaceXAIContext500KInput$2.00/MOutput$6.00/MProviders
+15
62.8±8.6
59.3±11.9E
59.7±10.8E
58.4±14.0P
Throughput
109 t/s
Latency
10.11s
Release date2026-08-12
  • 62.8
  • 59.3
  • 59.7
  • 58.4
DeepSeek V4 Pro 0813DeepSeekContext1.0MInputOutputProviders
57.5±7.7
57.1±6.1
61±8.3
58.4±18.0P
58.9±12.2E
54.9±16.0P
Throughput
Latency
Release date2026-08-12
  • 57.5
  • 57.1
  • 61
  • 58.4
  • 58.9
  • 54.9
Nemotron 3.5 LightningNVIDIAContext262.1KInput$0.070/MOutput$0.220/MProviders
+3
46.4±16.0P
49±13.9P
Throughput
213 t/s
Latency
21.68s
Release date2026-08-11
  • 46.4
  • 49
Solar Pro 4UpstageContext524.3KInput$0.300/MOutput$1.20/MProviders
54.8±16.0P
58±13.9P
Throughput
Latency
Release date2026-08-10
  • 54.8
  • 58
Muse Glimmer 30BMetaContext131.1KInputOutputProviders
47.7±10.2E
46.6±6.9
56.1±10.8E
43.3±16.0P
46.1±16.3P
46.4±9.4E
55.1±16.0P
Throughput
Latency
Release date2026-08-09
  • 47.7
  • 46.6
  • 56.1
  • 43.3
  • 46.1
  • 46.4
  • 55.1
Ling 3.0 TinyinclusionAIContext262.1KInputOutputProviders
40.4±16.0P
48.3±13.9P
Throughput
Latency
Release date2026-08-06
  • 40.4
  • 48.3
Muse Spark 1.2MetaContext1.0MInput$1.25/MOutput$4.25/MProviders
+2
56.9±9.3E
57.5±11.9E
61.2±10.8E
55.4±14.0P
Throughput
69 t/s
Latency
22.53s
Release date2026-08-05
  • 56.9
  • 57.5
  • 61.2
  • 55.4
Qwen3.8 MaxQwenContext1MInput$2.00/MOutput$6.00/MProviders
+22
66.4±10.9E
66.1±11.2E
64.4±10.1E
52.7±14.0P
64.7±6.3
57.7±16.0P
Throughput
87 t/s
Latency
3.46s
Release date2026-08-03
  • 66.4
  • 66.1
  • 64.4
  • 52.7
  • 64.7
  • 57.7
Inkling SmallThinking MachinesContext1.0MInput$0.300/MOutput$1.20/MProviders
55.7±10.9E
52.7±6.9
50.9±8.7
52.3±14.0P
52.5±14.4P
45.9±11.9E
57.4±16.0P
Throughput
Latency
Release date2026-07-30
  • 55.7
  • 52.7
  • 50.9
  • 52.3
  • 52.5
  • 45.9
  • 57.4
DeepSeek V4 Flash 0731DeepSeekContext1.3MInputOutputProviders
52.1±8.4
52.9±6.3
59.4±8.3
47.5±18.0P
57.9±12.2E
Throughput
74 t/s
Latency
2.70s
Release date2026-07-31
  • 52.1
  • 52.9
  • 59.4
  • 47.5
  • 57.9
Qwen3.7 FlashQwenContext1MInputOutputProviders
+4
Throughput
388 t/s
Latency
11.58s
Release date2026-07-27
Claude Opus 5AnthropicContext1MInput$5.00/MOutput$25.00/MProviders
+83
66.1±8.2
67.7±6.6
60.3±8.7
67.2±14.0P
64.3±17.1P
57.9±17.3P
Throughput
320 t/s
Latency
4.96s
Release date2026-07-24
  • 66.1
  • 67.7
  • 60.3
  • 67.2
  • 64.3
  • 57.9
Ling-3.0-flashInclusionaiContext262.1KInput$0.075/MOutput$0.220/MProviders
+10
47.8±10.2E
50.1±9.3E
52.7±10.8E
38.2±14.0P
46.7±11.5E
54.1±16.0P
Throughput
Latency
Release date2026-07-23
  • 47.8
  • 50.1
  • 52.7
  • 38.2
  • 46.7
  • 54.1
Laguna S 2.1PoolsideContext1.0MInputOutputProviders
+12
53.5±16.0P
54.1±9.3E
Throughput
41 t/s
Latency
1.46s
Release date2026-07-21
  • 53.5
  • 54.1
Gemini 3.5 Flash-LiteGoogleContext1.0MInput$0.300/MOutput$2.50/MProviders
+38
43.6±8.8
47.2±9.3E
53±10.8E
51.4±14.0P
45.8±17.3P
Throughput
Latency
Release date2026-07-21
  • 43.6
  • 47.2
  • 53
  • 51.4
  • 45.8
Gemini 3.6 FlashGoogleContext1.0MInput$0.750/MOutput$3.75/MProviders
+65
53.3±8.6
55.6±11.9E
59.7±10.8E
57.3±14.0P
50.6±17.3P
Throughput
737 t/s
Latency
2.34s
Release date2026-07-21
  • 53.3
  • 55.6
  • 59.7
  • 57.3
  • 50.6
InklingThinking MachinesContext1.0MInput$1.00/MOutput$4.05/MProviders
+12
50.9±8.3
49.8±6.9
55.4±10.8E
53.1±14.0P
65.4±16.3P
45.8±11.9E
56.3±16.0P
Throughput
25 t/s
Latency
0.79s
Release date2026-07-17
  • 50.9
  • 49.8
  • 55.4
  • 53.1
  • 65.4
  • 45.8
  • 56.3
Muse Spark 1.1MetaContext1.0MInput$1.25/MOutput$4.25/MProviders
+2
63.1±7.4
59±9.3E
59.5±10.2E
65.7±14.0P
56.8±16.1P
Throughput
Latency
Release date2026-07-16
  • 63.1
  • 59
  • 59.5
  • 65.7
  • 56.8
Kimi K3MoonshotAIContext1.0MInput$3.00/MOutput$15.00/MProviders
+82
68.3±7.3
57.3±11.9E
59.3±10.8E
61.4±14.0P
65±11.3E
Throughput
200 t/s
Latency
33.73s
Release date2026-07-16
  • 68.3
  • 57.3
  • 59.3
  • 61.4
  • 65
KAT-Coder-Pro V2.5KwaipilotContext256KInputOutputProviders
+2
Throughput
Latency
Release date2026-07-10
GPT-5.6 Sol ProOpenAIContext1.1MInputOutputProviders
+2
Throughput
Latency
Release date2026-07-09
GPT-5.6 Terra ProOpenAIContext1.1MInputOutputProviders
+2
Throughput
Latency
Release date2026-07-09
GPT-5.6 Luna ProOpenAIContext1.1MInputOutputProviders
+3
Throughput
Latency
Release date2026-07-09
GPT-3.5 Turbo 16kOpenAIContext16.4KInputOutputProviders
+1
Throughput
Latency
Release date2023-08-28
GPT-4 Turbo PreviewOpenAIContext128KInputOutputProviders
Throughput
Latency
Release date2024-01-25
GPT-3.5 Turbo (older v0613)OpenAIContext4.1KInputOutputProviders
+1
Throughput
Latency
Release date2024-01-25
Llama 3 8B InstructMetaContext8.2KInputOutputProviders
+1
Throughput
Latency
Release date
GPT-4o (2024-05-13)OpenAIContext128KInput$5.00/MOutput$15.00/MProviders
43.5±13.9P
42.7±11.1E
46±11.8E
Throughput
Latency
Release date2024-05-13
  • 43.5
  • 42.7
  • 46
GPT-4o-mini (2024-07-18)OpenAIContext128KInputOutputProviders
Throughput
Latency
Release date2024-07-18
Llama 3.1 70B InstructMetaContext131.1KInputOutputProviders
+1
Throughput
Latency
Release date2024-07-23
Llama 3.1 8B InstructMetaContext131.1KInputOutputProviders
+2
Throughput
Latency
Release date2024-07-23
GPT-4o (2024-08-06)OpenAIContext128KInput$2.50/MOutput$10.00/MProviders
44.3±13.9P
39.2±13.9P
46.2±11.8E
Throughput
Latency
Release date2024-08-06
  • 44.3
  • 39.2
  • 46.2
Hermes 3 70B InstructNousContext131.1KInput$0.700/MOutput$0.700/MProviders
36.6±13.9P
37.8±11.1E
39.6±11.8E
Throughput
Latency
Release date2024-08-18
  • 36.6
  • 37.8
  • 39.6
Llama 3.2 11B Vision InstructMetaContext131.1KInputOutputProviders
+4
Throughput
Latency
Release date2024-09-25
Llama 3.2 1B InstructMetaContext60KInputOutputProviders
+4
Throughput
Latency
Release date2024-09-25
Llama 3.2 3B InstructMetaContext131.1KInputOutputProviders
+8
Throughput
Latency
Release date2024-09-25
Mistral Large 2407MistralaiContext131.1KInput$2.00/MOutput$6.00/MProviders
40.4±13.9P
41.1±11.1E
44.5±11.8E
Throughput
Latency
Release date2024-11-19
  • 40.4
  • 41.1
  • 44.5
GPT-4o (2024-11-20)OpenAIContext128KInputOutputProviders
Throughput
Latency
Release date2024-11-20
Llama 3.3 70B InstructMetaContext131.1KInputOutputProviders
+6
Throughput
Latency
Release date2024-12-06
DeepSeek V3DeepSeekContext163.8KInputOutputProviders
+15
Throughput
Latency
Release date2024-12-26
R1 Distill Llama 70BDeepSeekContext8.2KInput$0.700/MOutput$1.10/MProviders
42.6±13.9P
45.3±11.1E
55±11.8E
Throughput
Latency
Release date2025-01-23
  • 42.6
  • 45.3
  • 55
Qwen2.5 VL 72B InstructQwenContext128KInputOutputProviders
Throughput
Latency
Release date2025-02-01
o3 Mini HighOpenAIContext200KInput$1.10/MOutput$4.40/MProviders
53.3±13.9P
51±11.1E
61.3±11.8E
Throughput
Latency
Release date2025-02-12
  • 53.3
  • 51
  • 61.3
Gemma 3 27BGoogleContext262.1KInputOutputProviders
Throughput
Latency
Release date2025-03-12
GPT-4o Search PreviewOpenAIContext128KInputOutputProviders
+2
Throughput
Latency
Release date2025-03-12
GPT-4o-mini Search PreviewOpenAIContext128KInputOutputProviders
+2
Throughput
Latency
Release date2025-03-12
Gemma 3 12BGoogleContext131.1KInputOutputProviders
Throughput
Latency
Release date2025-03-13
Gemma 3 4BGoogleContext131.1KInputOutputProviders
Throughput
Latency
Release date2025-03-13
Mistral Small 3.1 24BMistralContext128KInputOutputProviders
Throughput
Latency
Release date2025-03-17
o4 Mini HighOpenAIContext200KInputOutputProviders
51.7±16.1P
Throughput
Latency
Release date2025-04-16
  • 51.7
Qwen3 235B A22BQwenContext131.1KInput$0.700/MOutput$8.40/MProviders
43.1±13.9P
45.7±11.1E
51.1±11.8E
Throughput
Latency
Release date2025-04-28
  • 43.1
  • 45.7
  • 51.1
Qwen3 32BQwenContext131.1KInput$0.160/MOutput$0.640/MProviders
41.3±13.9P
46.4±11.1E
49.9±11.8E
Throughput
Latency
Release date2025-04-28
  • 41.3
  • 46.4
  • 49.9
Qwen3 14BQwenContext131.1KInput$0.350/MOutput$4.20/MProviders
40.3±13.9P
44.8±11.1E
49.8±11.8E
Throughput
Latency
Release date2025-04-28
  • 40.3
  • 44.8
  • 49.8
Qwen3 8BQwenContext131.1KInput$0.180/MOutput$2.10/MProviders
32.6±13.9P
42.2±11.1E
48.5±11.8E
Throughput
Latency
Release date2025-04-28
  • 32.6
  • 42.2
  • 48.5
Qwen3 30B A3BQwenContext131.1KInput$0.200/MOutput$2.40/MProviders
40.9±13.9P
43.2±11.1E
49.4±11.8E
Throughput
Latency
Release date2025-04-28
  • 40.9
  • 43.2
  • 49.4
Llama Guard 4 12BMetaContext1.0MInputOutputProviders
Throughput
Latency
Release date2025-04-30