ModelContextInputOutputProvidersAgentsCodingReasoningKnowledgeMathMultilingualMultimodalInstruction followingThroughputLatencyRelease date
DeepSeek V3ContextInput$0.360/MOutput$0.890/MProviders
+150
34.6±11.8E
37.4±11.2E
44.4±8.7
41.1±16.0P
49.5±9.3E
39.7±11.9E
Throughput
49 t/s
Latency
3.39s
Release date
  • 34.6
  • 37.4
  • 44.4
  • 41.1
  • 49.5
  • 39.7
Qwen3Context262.1KInput$0.200/MOutput$0.800/MProviders
+126
45.7±13.9P
47.6±11.1E
36.9±16.3P
58.3±11.8E
36.7±16.4P
Throughput
100 t/s
Latency
12.55s
Release date2025-07-21
  • 45.7
  • 47.6
  • 36.9
  • 58.3
  • 36.7
Qwen3 VLContextInputOutputProviders
+17
Throughput
206 t/s
Latency
6.63s
Release date
Devstral 2ContextInputOutputProviders
+4
46.2±13.9P
45.3±11.1E
43.2±16.1P
Throughput
71 t/s
Latency
1.34s
Release date
  • 46.2
  • 45.3
  • 43.2
Ministral 3ContextInput$0.200/MOutput$0.200/MProviders
+2
31.5±13.9P
36.6±11.1E
40.5±16.1P
Throughput
208 t/s
Latency
2.84s
Release date
  • 31.5
  • 36.6
  • 40.5
Mistral Large 3ContextInput$0.500/MOutput$1.50/MProviders
+7
39.3±11.8E
47.9±13.9P
44.5±8.7
41.6±16.0P
43.5±16.1P
42.2±16.0P
Throughput
29 t/s
Latency
3.00s
Release date
  • 39.3
  • 47.9
  • 44.5
  • 41.6
  • 43.5
  • 42.2
Qwen3 CoderQwenContext262.1KInputOutputProviders
+61
Throughput
243 t/s
Latency
5.28s
Release date2025-07-23
Qwen3 EmbeddingContextInputOutputProviders
+42
Throughput
Latency
Release date
Qwen3 NextContextInputOutputProviders
+7
Throughput
Latency
Release date
DeepSeek CoderContextInputOutputProviders
+8
Throughput
Latency
Release date
Qwen ImageContextInputOutputProviders
+34
Throughput
Latency
Release date
Phi 4MicrosoftContext16.4KInput$0.125/MOutput$0.500/MProviders
+17
28.4±16.0P
39.2±13.9P
39.1±8.7
37.5±16.0P
46.9±11.8E
37.8±16.0P
Throughput
43 t/s
Latency
1.32s
Release date2025-01-10
  • 28.4
  • 39.2
  • 39.1
  • 37.5
  • 46.9
  • 37.8
Intern-S1ContextInputOutputProviders
+4
Throughput
Latency
Release date
MiniMax M1MiniMaxContext1MInput$0.550/MOutput$2.20/MProviders
+30
51.2±13.9P
49.4±11.1E
58.9±11.8E
Throughput
Latency
Release date2025-06-17
  • 51.2
  • 49.4
  • 58.9
MAI-DS-R1ContextInputOutputProviders
+10
Throughput
Latency
Release date
DeepSeek VL2ContextInputOutputProviders
+8
Throughput
126 t/s
Latency
0.75s
Release date
GLM-4ContextInputOutputProviders
+56
Throughput
61 t/s
Latency
0.87s
Release date
GPT-OSSContext131.1KInputOutputProviders
+137
Throughput
383 t/s
Latency
2.36s
Release date2025-08-05
GPT-4oOpenAIContext128KInput$2.50/MOutput$10.00/MProviders
+132
39±16.0P
46±13.9P
39.9±8.7
49±16.0P
42.2±9.3E
41.6±16.0P
Throughput
82 t/s
Latency
3.44s
Release date2024-11-20
  • 39
  • 46
  • 39.9
  • 49
  • 42.2
  • 41.6
MiniMax M2.5MiniMaxContext204.8KInput$0.300/MOutput$1.20/MProviders
+186
50.2±11.8E
54.6±13.9P
Throughput
61 t/s
Latency
9.68s
Release date2026-02-12
  • 50.2
  • 54.6
GPT-5.4 MiniOpenAIContext400KInput$0.750/MOutput$4.50/MProviders
+238
51±8.1
57.5±11.8E
55.6±10.8E
46.6±14.0P
49.7±16.1P
47.1±16.1P
53.6±16.0P
Throughput
134 t/s
Latency
3.54s
Release date2026-03-17
  • 51
  • 57.5
  • 55.6
  • 46.6
  • 49.7
  • 47.1
  • 53.6
MiniMax Hailuo 2.3ContextInputOutputProviders
+16
Throughput
Latency
Release date
GLM-5Z.aiContext204.8KInput$1.00/MOutput$3.20/MProviders
+185
41.1±5.6
51.5±8.1
48.9±8.1
43.3±16.3P
51.1±9.1E
43.8±12.2E
52.4±11.9E
Throughput
51 t/s
Latency
21.59s
Release date2026-02-11
  • 41.1
  • 51.5
  • 48.9
  • 43.3
  • 51.1
  • 43.8
  • 52.4
GPT-4o MiniOpenAIContext128KInput$0.150/MOutput$0.600/MProviders
+121
30.8±16.0P
37.4±13.9P
40.3±11.1E
53.4±16.0P
46.1±11.8E
40.5±16.0P
Throughput
84 t/s
Latency
4.10s
Release date2024-07-18
  • 30.8
  • 37.4
  • 40.3
  • 53.4
  • 46.1
  • 40.5
Claude Sonnet 4.5AnthropicContext1MInput$3.00/MOutput$15.00/MProviders
+174
46.6±8.7
51.8±11.2E
45.6±8.9E
54.1±16.1P
42.5±12.3E
Throughput
41 t/s
Latency
4.57s
Release date2025-09-29
  • 46.6
  • 51.8
  • 45.6
  • 54.1
  • 42.5
GPT-5.1OpenAIContext400KInput$1.25/MOutput$10.00/MProviders
+153
47.7±9.3E
49.1±11.2E
54±8.7
53±16.0P
54±16.1P
53.5±16.0P
Throughput
142 t/s
Latency
2.78s
Release date2025-11-13
  • 47.7
  • 49.1
  • 54
  • 53
  • 54
  • 53.5
GLM-4.6VZ.aiContext131.1KInput$0.300/MOutput$0.900/MProviders
+60
40.1±13.9P
49.8±11.1E
52.2±16.1P
Throughput
40 t/s
Latency
27.37s
Release date2025-12-08
  • 40.1
  • 49.8
  • 52.2
Gemini 2.5 ProGoogleContext1.0MInput$1.25/MOutput$10.00/MProviders
+171
43.6±9.3E
39.1±8.9
54.2±8.7
35±14.0P
56.1±9.3E
45.9±16.0P
Throughput
88 t/s
Latency
16.94s
Release date2025-06-17
  • 43.6
  • 39.1
  • 54.2
  • 35
  • 56.1
  • 45.9
O1OpenAIContext200KInput$15.00/MOutput$60.00/MProviders
+75
45.6±16.0P
50.6±13.9P
50.7±8.7
62.5±17.4P
52.6±9.3E
51.5±11.9E
Throughput
Latency
Release date2024-12-17
  • 45.6
  • 50.6
  • 50.7
  • 62.5
  • 52.6
  • 51.5
O1 MiniOpenAIContextInputOutputProviders
+60
47.4±13.9P
44.7±11.1E
54.9±11.8E
Throughput
32 t/s
Latency
13.84s
Release date
  • 47.4
  • 44.7
  • 54.9
GPT-4.1OpenAIContext1.0MInput$2.00/MOutput$8.00/MProviders
+109
40±11.8E
38.3±11.2E
48.6±8.7
57±17.4P
49±9.3E
42.1±11.9E
Throughput
85 t/s
Latency
1.92s
Release date2025-04-14
  • 40
  • 38.3
  • 48.6
  • 57
  • 49
  • 42.1
O3OpenAIContext200KInput$2.00/MOutput$8.00/MProviders
+85
49.3±16.0P
55.2±13.9P
54.3±8.7
47.7±16.0P
58.8±9.3E
53±16.0P
Throughput
134 t/s
Latency
2.82s
Release date2025-04-16
  • 49.3
  • 55.2
  • 54.3
  • 47.7
  • 58.8
  • 53
Kimi K2MoonshotAIContext131.1KInput$0.570/MOutput$2.30/MProviders
+87
45.3±16.0P
48.3±13.9P
48.6±8.7
44.5±16.0P
53.7±9.3E
43.8±16.0P
Throughput
29 t/s
Latency
2.21s
Release date2025-09-04
  • 45.3
  • 48.3
  • 48.6
  • 44.5
  • 53.7
  • 43.8
Claude Opus 4.6AnthropicContext1MInput$5.00/MOutput$25.00/MProviders
+228
54.9±5.7
58.9±8.0
53.4±8.7
53.7±13.7P
56.5±16.1P
54±17.3P
48.4±11.1E
44.7±16.0P
Throughput
45 t/s
Latency
5.30s
Release date2026-02-04
  • 54.9
  • 58.9
  • 53.4
  • 53.7
  • 56.5
  • 54
  • 48.4
  • 44.7
Gemini 3 ProGoogleContextInput$2.00/MOutput$12.00/MProviders
+139
49±9.3E
55.5±11.2E
54.1±6.4
56.1±14.0P
55.7±16.1P
54±17.3P
49.9±6.5
52.6±16.0P
Throughput
73 t/s
Latency
10.59s
Release date
  • 49
  • 55.5
  • 54.1
  • 56.1
  • 55.7
  • 54
  • 49.9
  • 52.6
Phi 4 Mini InstructContext131.1KInputOutputProviders
+28
27.9±13.9P
34.3±11.1E
42±11.8E
Throughput
Latency
Release date2025-10-17
  • 27.9
  • 34.3
  • 42
GPT-5 MiniOpenAIContext400KInput$0.250/MOutput$2.00/MProviders
+111
49.5±11.2E
53.5±11.1E
52.2±16.1P
Throughput
90 t/s
Latency
7.03s
Release date2025-08-07
  • 49.5
  • 53.5
  • 52.2
GPT-5.1 Codex MaxOpenAIContext400KInputOutputProviders
+126
49.9±16.0P
50.8±11.8E
54.3±10.8E
50±16.0P
52.5±16.0P
Throughput
56 t/s
Latency
6.29s
Release date2025-12-04
  • 49.9
  • 50.8
  • 54.3
  • 50
  • 52.5
Gemini 1.5 ProGoogleContextInputOutputProviders
+21
40.2±13.9P
39.7±11.1E
54.2±16.0P
43.7±11.8E
Throughput
18 t/s
Latency
2.58s
Release date
  • 40.2
  • 39.7
  • 54.2
  • 43.7
GPT-4.1 MiniOpenAIContext1.0MInput$0.400/MOutput$1.60/MProviders
+112
40.3±11.8E
39±11.2E
45.4±8.7
49.3±17.4P
49.9±9.3E
42.4±11.9E
Throughput
93 t/s
Latency
2.69s
Release date2025-04-14
  • 40.3
  • 39
  • 45.4
  • 49.3
  • 49.9
  • 42.4
Grok Imagine VideoSpaceXAIContext0InputOutputProviders
+21
Throughput
Latency
Release date2026-05-18
Claude Opus 4AnthropicContext200KInput$15.00/MOutput$75.00/MProviders
+91
51±13.9P
51.7±11.1E
54.4±11.8E
Throughput
25 t/s
Latency
12.05s
Release date2025-05-22
  • 51
  • 51.7
  • 54.4
Gemini 2.5 Flash ImageGoogleContext32.8KInputOutputProviders
+93
Throughput
Latency
Release date2025-10-07
GPT-5.4OpenAIContext1.1MInput$2.50/MOutput$15.00/MProviders
+292
57.2±5.1
62.4±10.3E
56.2±8.7
59.9±15.1P
57.6±16.1P
52.9±6.6
53.9±16.0P
Throughput
49 t/s
Latency
4.45s
Release date2026-03-05
  • 57.2
  • 62.4
  • 56.2
  • 59.9
  • 57.6
  • 52.9
  • 53.9
GPT Audio MiniOpenAIContext128KInputOutputProviders
+16
Throughput
Latency
Release date2026-01-19
GLM-4.7Z.aiContext204.8KInput$0.600/MOutput$2.20/MProviders
+151
46.7±8.6
53±10.5E
54.4±8.7
34.7±14.0P
48.8±12.3E
51.8±16.0P
Throughput
94 t/s
Latency
18.83s
Release date2025-12-22
  • 46.7
  • 53
  • 54.4
  • 34.7
  • 48.8
  • 51.8
Claude Opus 4.5AnthropicContext200KInput$5.00/MOutput$25.00/MProviders
+166
48.5±5.3
55.7±8.1
60.6±8.1
57.6±11.6E
49.9±9.1E
51.7±12.2E
32.4±6.5
45.7±11.9E
Throughput
54 t/s
Latency
3.12s
Release date2025-11-24
  • 48.5
  • 55.7
  • 60.6
  • 57.6
  • 49.9
  • 51.7
  • 32.4
  • 45.7
GLM-4.6Z.aiContext204.8KInput$0.550/MOutput$2.20/MProviders
+101
48.4±16.0P
41.9±11.2E
42.9±8.7
43.6±16.0P
44.8±16.1P
42.4±16.0P
Throughput
53 t/s
Latency
16.68s
Release date2025-09-30
  • 48.4
  • 41.9
  • 42.9
  • 43.6
  • 44.8
  • 42.4
Gemini 2.5 Flash LiteGoogleContext1.0MInput$0.100/MOutput$0.400/MProviders
+128
36.5±13.9P
42.9±11.1E
53.3±11.8E
Throughput
277 t/s
Latency
1.41s
Release date2025-09-25
  • 36.5
  • 42.9
  • 53.3
DeepSeek R1DeepSeekContext64KInput$1.35/MOutput$3.00/MProviders
+149
41.2±16.0P
49.7±13.9P
50.2±8.7
44.7±16.0P
57±11.8E
43.3±16.0P
Throughput
52 t/s
Latency
10.36s
Release date2025-05-28
  • 41.2
  • 49.7
  • 50.2
  • 44.7
  • 57
  • 43.3
Devstral SmallMistral AIContextInputOutputProviders
+7
38.8±13.9P
39.5±11.1E
43.4±11.8E
Throughput
Latency
Release date
  • 38.8
  • 39.5
  • 43.4
GPT-3.5 Turbo InstructOpenAIContext4.1KInputOutputProviders
+44
Throughput
Latency
Release date2023-09-28
GPT-3.5 TurboOpenAIContext16.4KInput$0.500/MOutput$1.50/MProviders
+98
48.8±18.4P
33.5±11.8E
37.8±16.0P
Throughput
115 t/s
Latency
2.21s
Release date2024-01-25
  • 48.8
  • 33.5
  • 37.8
Gemini 3.1 Flash LiteGoogleContext1.0MInputOutputProviders
+128
44.5±16.1P
27.5±16.1P
42.6±16.1P
Throughput
191 t/s
Latency
6.58s
Release date2026-05-07
  • 44.5
  • 27.5
  • 42.6
O3 ProOpenAIContext200KInput$20.00/MOutput$80.00/MProviders
+61
54±16.0P
60.1±16.0P
Throughput
Latency
Release date2025-06-10
  • 54
  • 60.1
GPT-4 TurboOpenAIContext128KInput$10.00/MOutput$30.00/MProviders
+73
43.3±13.9P
43.2±11.8E
53.6±16.0P
45.8±11.8E
Throughput
50 t/s
Latency
0.97s
Release date2024-04-09
  • 43.3
  • 43.2
  • 53.6
  • 45.8
GPT-5.1 Codex MiniOpenAIContext400KInput$0.250/MOutput$2.00/MProviders
+125
56.6±13.9P
53.2±11.1E
53.4±16.1P
Throughput
136 t/s
Latency
4.28s
Release date2025-11-13
  • 56.6
  • 53.2
  • 53.4
Claude Opus 4.1AnthropicContext200KInput$15.00/MOutput$75.00/MProviders
+100
46.5±16.2P
47.1±16.1P
59±16.0P
Throughput
22 t/s
Latency
3.56s
Release date2025-08-05
  • 46.5
  • 47.1
  • 59
Qwen3 Omni FlashContextInputOutputProviders
+38
Throughput
62 t/s
Latency
5.13s
Release date
DeepSeek ReasonerDeepSeekContextInputOutputProviders
+85
Throughput
28 t/s
Latency
6.40s
Release date