AI Model Benchmark Directory

LMSpeed is an AI model directory for comparing API price, output speed, first-token latency, provider coverage, and benchmark data. Use it to narrow your model and provider choice, then test the endpoint that fits your workload.

Price, speed, latency, and availability can change. Treat this table as a current signal and verify your own endpoint before you deploy.

ModelContextInputOutputProvidersAgentsCodingReasoningKnowledgeMathMultilingualMultimodalInstruction followingThroughputLatencyRelease date
DeepSeek V3ContextInput$0.360/MOutput$0.890/MProviders
+151
35.2±11.8E
37.3±11.2E
44.3±8.7
40.9±16.0P
49.5±9.3E
39.7±11.9E
Throughput
49 t/s
Latency
3.39s
Release date
  • 35.2
  • 37.3
  • 44.3
  • 40.9
  • 49.5
  • 39.7
Qwen3Context262.1KInput$0.200/MOutput$0.800/MProviders
+126
45.7±13.9P
47.6±11.1E
36.9±16.3P
58.3±11.8E
36.7±16.4P
Throughput
100 t/s
Latency
12.55s
Release date2025-07-21
  • 45.7
  • 47.6
  • 36.9
  • 58.3
  • 36.7
Qwen3 VLContextInputOutputProviders
+17
Throughput
206 t/s
Latency
6.63s
Release date
Devstral 2ContextInputOutputProviders
+4
46.1±13.9P
45.2±11.1E
43.2±16.1P
Throughput
71 t/s
Latency
1.34s
Release date
  • 46.1
  • 45.2
  • 43.2
Ministral 3ContextInput$0.200/MOutput$0.200/MProviders
+2
31.8±13.9P
36.5±11.1E
40.5±16.1P
Throughput
208 t/s
Latency
2.84s
Release date
  • 31.8
  • 36.5
  • 40.5
Mistral Large 3ContextInput$0.500/MOutput$1.50/MProviders
+7
39.6±11.8E
47.8±13.9P
44.5±8.7
41.5±16.0P
43.5±16.1P
42.2±16.0P
Throughput
29 t/s
Latency
3.00s
Release date
  • 39.6
  • 47.8
  • 44.5
  • 41.5
  • 43.5
  • 42.2
Qwen3 CoderQwenContext262.1KInputOutputProviders
+61
Throughput
243 t/s
Latency
5.28s
Release date2025-07-23
Qwen3 EmbeddingContextInputOutputProviders
+43
Throughput
Latency
Release date
Qwen3 NextContextInputOutputProviders
+7
Throughput
Latency
Release date
DeepSeek CoderContextInputOutputProviders
+8
Throughput
Latency
Release date
Qwen ImageContextInputOutputProviders
+34
Throughput
Latency
Release date
Phi 4MicrosoftContext16.4KInput$0.125/MOutput$0.500/MProviders
+17
28.3±16.0P
39.3±13.9P
39.1±8.7
37.4±16.0P
46.9±11.8E
37.8±16.0P
Throughput
43 t/s
Latency
1.32s
Release date2025-01-10
  • 28.3
  • 39.3
  • 39.1
  • 37.4
  • 46.9
  • 37.8
Intern-S1ContextInputOutputProviders
+4
Throughput
Latency
Release date
MiniMax M1MiniMaxContext1MInput$0.550/MOutput$2.20/MProviders
+31
51.1±13.9P
49.3±11.1E
58.9±11.8E
Throughput
Latency
Release date2025-06-17
  • 51.1
  • 49.3
  • 58.9
MAI-DS-R1ContextInputOutputProviders
+10
Throughput
Latency
Release date
DeepSeek VL2ContextInputOutputProviders
+8
Throughput
126 t/s
Latency
0.75s
Release date
GLM-4ContextInputOutputProviders
+57
Throughput
61 t/s
Latency
0.87s
Release date
GPT-OSSContext131.1KInputOutputProviders
+137
Throughput
374 t/s
Latency
3.23s
Release date2025-08-05
GPT-4oOpenAIContext128KInput$2.50/MOutput$10.00/MProviders
+133
39±16.0P
46±13.9P
39.9±8.7
48.9±16.0P
42.2±9.3E
41.6±16.0P
Throughput
82 t/s
Latency
3.44s
Release date2024-11-20
  • 39
  • 46
  • 39.9
  • 48.9
  • 42.2
  • 41.6
MiniMax M2.5MiniMaxContext204.8KInput$0.300/MOutput$1.20/MProviders
+188
50±11.8E
54.5±13.9P
Throughput
61 t/s
Latency
9.68s
Release date2026-02-12
  • 50
  • 54.5
GPT-5.4 MiniOpenAIContext400KInput$0.750/MOutput$4.50/MProviders
+240
50.9±8.1
57.2±11.8E
55.5±10.8E
46.1±14.0P
49.7±16.1P
47.1±16.1P
53.6±16.0P
Throughput
134 t/s
Latency
3.54s
Release date2026-03-17
  • 50.9
  • 57.2
  • 55.5
  • 46.1
  • 49.7
  • 47.1
  • 53.6
MiniMax Hailuo 2.3ContextInputOutputProviders
+16
Throughput
Latency
Release date
GLM-5Z.aiContext204.8KInput$1.00/MOutput$3.20/MProviders
+188
41±5.6
51.1±8.1
48.8±8.1
43.3±16.3P
51.1±9.1E
43.8±12.2E
52.4±11.9E
Throughput
51 t/s
Latency
21.59s
Release date2026-02-11
  • 41
  • 51.1
  • 48.8
  • 43.3
  • 51.1
  • 43.8
  • 52.4
GPT-4o MiniOpenAIContext128KInput$0.150/MOutput$0.600/MProviders
+121
32±16.0P
37.5±13.9P
40.2±11.1E
53.2±16.0P
46.1±11.8E
40.5±16.0P
Throughput
84 t/s
Latency
4.10s
Release date2024-07-18
  • 32
  • 37.5
  • 40.2
  • 53.2
  • 46.1
  • 40.5
Claude Sonnet 4.5AnthropicContext1MInput$3.00/MOutput$15.00/MProviders
+174
46.5±8.7
51.6±11.2E
45.5±8.9E
54.2±16.1P
42.5±12.3E
Throughput
41 t/s
Latency
4.57s
Release date2025-09-29
  • 46.5
  • 51.6
  • 45.5
  • 54.2
  • 42.5
GPT-5.1OpenAIContext400KInput$1.25/MOutput$10.00/MProviders
+153
47.7±9.3E
49.1±11.2E
53.9±8.7
52.9±16.0P
54±16.1P
53.5±16.0P
Throughput
142 t/s
Latency
2.78s
Release date2025-11-13
  • 47.7
  • 49.1
  • 53.9
  • 52.9
  • 54
  • 53.5
GLM-4.6VZ.aiContext131.1KInput$0.300/MOutput$0.900/MProviders
+61
40.1±13.9P
49.7±11.1E
52.2±16.1P
Throughput
40 t/s
Latency
27.37s
Release date2025-12-08
  • 40.1
  • 49.7
  • 52.2
Gemini 2.5 ProGoogleContext1.0MInput$1.25/MOutput$10.00/MProviders
+172
43.8±9.3E
38.9±8.9
54.1±8.7
36.3±14.0P
56.1±9.3E
45.9±16.0P
Throughput
88 t/s
Latency
16.94s
Release date2025-06-17
  • 43.8
  • 38.9
  • 54.1
  • 36.3
  • 56.1
  • 45.9
O1OpenAIContext200KInput$15.00/MOutput$60.00/MProviders
+75
45.5±16.0P
50.5±13.9P
50.7±8.7
62.5±17.4P
52.6±9.3E
51.5±11.9E
Throughput
Latency
Release date2024-12-17
  • 45.5
  • 50.5
  • 50.7
  • 62.5
  • 52.6
  • 51.5
O1 MiniOpenAIContextInputOutputProviders
+60
47.4±13.9P
44.6±11.1E
54.9±11.8E
Throughput
32 t/s
Latency
13.84s
Release date
  • 47.4
  • 44.6
  • 54.9
GPT-4.1OpenAIContext1.0MInput$2.00/MOutput$8.00/MProviders
+109
40±11.8E
38.1±11.2E
48.5±8.7
57±17.4P
49±9.3E
42.1±11.9E
Throughput
85 t/s
Latency
1.92s
Release date2025-04-14
  • 40
  • 38.1
  • 48.5
  • 57
  • 49
  • 42.1
O3OpenAIContext200KInput$2.00/MOutput$8.00/MProviders
+85
49.3±16.0P
55.1±13.9P
54.3±8.7
47.6±16.0P
58.8±9.3E
53±16.0P
Throughput
134 t/s
Latency
2.82s
Release date2025-04-16
  • 49.3
  • 55.1
  • 54.3
  • 47.6
  • 58.8
  • 53
Kimi K2MoonshotAIContext131.1KInput$0.570/MOutput$2.30/MProviders
+87
45.3±16.0P
48.2±13.9P
48.5±8.7
44.3±16.0P
53.7±9.3E
43.8±16.0P
Throughput
29 t/s
Latency
2.21s
Release date2025-09-04
  • 45.3
  • 48.2
  • 48.5
  • 44.3
  • 53.7
  • 43.8
Claude Opus 4.6AnthropicContext1MInput$5.00/MOutput$25.00/MProviders
+229
54.8±5.7
58.6±8.0
53.3±8.7
53.7±13.7P
56.5±16.1P
54±17.3P
48.4±11.1E
44.7±16.0P
Throughput
45 t/s
Latency
5.30s
Release date2026-02-04
  • 54.8
  • 58.6
  • 53.3
  • 53.7
  • 56.5
  • 54
  • 48.4
  • 44.7
Gemini 3 ProGoogleContextInput$2.00/MOutput$12.00/MProviders
+139
49±9.3E
55.3±11.2E
53.9±6.4
55.1±14.0P
55.7±16.1P
54±17.3P
49.9±6.5
52.6±16.0P
Throughput
73 t/s
Latency
10.59s
Release date
  • 49
  • 55.3
  • 53.9
  • 55.1
  • 55.7
  • 54
  • 49.9
  • 52.6
Phi 4 Mini InstructContext131.1KInputOutputProviders
+28
27.8±13.9P
34.3±11.1E
42±11.8E
Throughput
Latency
Release date2025-10-17
  • 27.8
  • 34.3
  • 42
GPT-5 MiniOpenAIContext400KInput$0.250/MOutput$2.00/MProviders
+111
49.4±11.2E
53.5±11.1E
52.2±16.1P
Throughput
90 t/s
Latency
7.03s
Release date2025-08-07
  • 49.4
  • 53.5
  • 52.2
GPT-5.1 Codex MaxOpenAIContext400KInputOutputProviders
+126
49.9±16.0P
50.7±11.8E
54.2±10.8E
49.9±16.0P
52.5±16.0P
Throughput
56 t/s
Latency
6.29s
Release date2025-12-04
  • 49.9
  • 50.7
  • 54.2
  • 49.9
  • 52.5
Gemini 1.5 ProGoogleContextInputOutputProviders
+21
40.3±13.9P
39.6±11.1E
54±16.0P
43.7±11.8E
Throughput
18 t/s
Latency
2.58s
Release date
  • 40.3
  • 39.6
  • 54
  • 43.7
GPT-4.1 MiniOpenAIContext1.0MInput$0.400/MOutput$1.60/MProviders
+112
40.7±11.8E
38.8±11.2E
45.3±8.7
49.3±17.4P
49.9±9.3E
42.4±11.9E
Throughput
93 t/s
Latency
2.69s
Release date2025-04-14
  • 40.7
  • 38.8
  • 45.3
  • 49.3
  • 49.9
  • 42.4
Grok Imagine VideoSpaceXAIContext0InputOutputProviders
+22
Throughput
Latency
Release date2026-05-18
Claude Opus 4AnthropicContext200KInput$15.00/MOutput$75.00/MProviders
+91
50.9±13.9P
51.6±11.1E
54.4±11.8E
Throughput
25 t/s
Latency
12.05s
Release date2025-05-22
  • 50.9
  • 51.6
  • 54.4
Gemini 2.5 Flash ImageGoogleContext32.8KInputOutputProviders
+93
Throughput
Latency
Release date2025-10-07
GPT-5.4OpenAIContext1.1MInput$2.50/MOutput$15.00/MProviders
+294
57±5.1
62.1±10.3E
55.9±8.7
59.9±15.1P
57.6±16.1P
52.9±6.6
53.9±16.0P
Throughput
49 t/s
Latency
4.45s
Release date2026-03-05
  • 57
  • 62.1
  • 55.9
  • 59.9
  • 57.6
  • 52.9
  • 53.9
GPT Audio MiniOpenAIContext128KInputOutputProviders
+16
Throughput
Latency
Release date2026-01-19
GLM-4.7Z.aiContext204.8KInput$0.600/MOutput$2.20/MProviders
+152
46.6±8.6
52.8±10.5E
54.4±8.7
34.1±14.0P
48.8±12.3E
51.8±16.0P
Throughput
95 t/s
Latency
19.09s
Release date2025-12-22
  • 46.6
  • 52.8
  • 54.4
  • 34.1
  • 48.8
  • 51.8
Claude Opus 4.5AnthropicContext200KInput$5.00/MOutput$25.00/MProviders
+166
48.4±5.3
55.2±8.1
60.6±8.1
57.6±11.6E
49.9±9.1E
51.7±12.2E
32.4±6.5
45.7±11.9E
Throughput
54 t/s
Latency
3.12s
Release date2025-11-24
  • 48.4
  • 55.2
  • 60.6
  • 57.6
  • 49.9
  • 51.7
  • 32.4
  • 45.7
GLM-4.6Z.aiContext204.8KInput$0.550/MOutput$2.20/MProviders
+102
48.3±16.0P
41.9±11.2E
42.8±8.7
43.5±16.0P
44.8±16.1P
42.4±16.0P
Throughput
53 t/s
Latency
16.68s
Release date2025-09-30
  • 48.3
  • 41.9
  • 42.8
  • 43.5
  • 44.8
  • 42.4
Gemini 2.5 Flash LiteGoogleContext1.0MInput$0.100/MOutput$0.400/MProviders
+128
36.7±13.9P
42.8±11.1E
53.3±11.8E
Throughput
277 t/s
Latency
1.41s
Release date2025-09-25
  • 36.7
  • 42.8
  • 53.3
DeepSeek R1DeepSeekContext64KInput$1.35/MOutput$3.00/MProviders
+149
41.2±16.0P
49.6±13.9P
50.1±8.7
44.6±16.0P
57±11.8E
43.3±16.0P
Throughput
52 t/s
Latency
10.36s
Release date2025-05-28
  • 41.2
  • 49.6
  • 50.1
  • 44.6
  • 57
  • 43.3
Devstral SmallMistral AIContextInputOutputProviders
+7
38.9±13.9P
39.4±11.1E
43.4±11.8E
Throughput
Latency
Release date
  • 38.9
  • 39.4
  • 43.4
GPT-3.5 Turbo InstructOpenAIContext4.1KInputOutputProviders
+44
Throughput
Latency
Release date2023-09-28
GPT-3.5 TurboOpenAIContext16.4KInput$0.500/MOutput$1.50/MProviders
+98
48.8±18.4P
33.3±11.8E
37.8±16.0P
Throughput
115 t/s
Latency
2.21s
Release date2024-01-25
  • 48.8
  • 33.3
  • 37.8
Gemini 3.1 Flash LiteGoogleContext1.0MInputOutputProviders
+128
44.5±16.1P
27.4±16.1P
42.6±16.1P
Throughput
191 t/s
Latency
6.58s
Release date2026-05-07
  • 44.5
  • 27.4
  • 42.6
O3 ProOpenAIContext200KInput$20.00/MOutput$80.00/MProviders
+61
54±16.0P
60±16.0P
Throughput
Latency
Release date2025-06-10
  • 54
  • 60
GPT-4 TurboOpenAIContext128KInput$10.00/MOutput$30.00/MProviders
+73
43.3±13.9P
43.2±11.8E
53.5±16.0P
45.8±11.8E
Throughput
50 t/s
Latency
0.97s
Release date2024-04-09
  • 43.3
  • 43.2
  • 53.5
  • 45.8
GPT-5.1 Codex MiniOpenAIContext400KInput$0.250/MOutput$2.00/MProviders
+125
56.4±13.9P
53.1±11.1E
53.4±16.1P
Throughput
136 t/s
Latency
4.28s
Release date2025-11-13
  • 56.4
  • 53.1
  • 53.4
Claude Opus 4.1AnthropicContext200KInput$15.00/MOutput$75.00/MProviders
+100
46.4±16.2P
46.9±16.1P
58.9±16.0P
Throughput
22 t/s
Latency
3.56s
Release date2025-08-05
  • 46.4
  • 46.9
  • 58.9
Qwen3 Omni FlashContextInputOutputProviders
+39
Throughput
62 t/s
Latency
5.13s
Release date
DeepSeek ReasonerDeepSeekContextInputOutputProviders
+85
Throughput
28 t/s
Latency
6.40s
Release date

How to read category scores

Agents, Coding, Reasoning, and the other capability columns are 0–100 observed-capability estimates relative to the eligible model population in a dated Category Score V3 run.

They are not success rates, IQ scores, or an average across all eight categories. Read them with the 80% interval, measured dimensions, benchmark families, and evidence shown on each model page.

Read the complete scoring methodology

What this directory shows

Each row brings together the information you need to compare an LLM API. Some fields are blank when LMSpeed has no current data for that model or provider.

API price
Compare input and output price per million tokens when it is available.
Speed and latency
Use throughput and first-token latency to compare response behavior.
Provider coverage
Open a model to review its listed providers and their current details.
Capability data
Use the capability columns when current model scores are available.

How to use the model directory

Start with the decision that matters most for your workload, then compare the current rows before you test an endpoint.

  1. Find a model.Search by model name, slug, or description.
  2. Sort the key metric.Sort by price, throughput, latency, provider count, or capability data.
  3. Compare providers.Open a model page to review the provider options and current data.
  4. Test your endpoint.Run a speed test before you use an endpoint in production.

Frequently Asked Questions

How does LMSpeed benchmark AI models?

LMSpeed runs standardized five-round API speed tests on each model, measuring output throughput (tokens per second), first-token latency, and total response time across multiple providers.

Which AI model has the lowest API latency?

Latency varies by provider and model. Use the LMSpeed model directory to sort by latency and find the model with the fastest first-token response time. Check the latency leaderboard for monthly rankings.

How to compare LLM API pricing across providers?

LMSpeed lists input and output token prices per million for each model across available providers. Sort by price to find a lower-cost option, or filter by provider to compare rates side by side.

What is an AI model directory?

An AI model directory is a searchable list of models and their comparison data. On LMSpeed, it brings together API price, throughput, first-token latency, provider coverage, and capability data when available.

How do I choose a model and provider?

Start with the metric that matters most for your workload. Then open a model page to compare provider data. Run your own speed test before you use an endpoint in production.

Can model price and speed change?

Yes. Price, availability, throughput, and latency can change by model, provider, and time. Use current page values as a comparison signal and confirm with your own endpoint test.