AI Model Benchmark Directory

LMSpeed is an AI model directory for comparing API price, output speed, first-token latency, provider coverage, and benchmark data. Use it to narrow your model and provider choice, then test the endpoint that fits your workload.

Price, speed, latency, and availability can change. Treat this table as a current signal and verify your own endpoint before you deploy.

ModelContextInputOutputProvidersAgentsCodingReasoningKnowledgeMathMultilingualMultimodalInstruction followingThroughputLatencyRelease date
Qwen3.8 Max (0902)QwenContext1MInputOutputProviders
+1
Throughput
Latency
Release date2026-09-03
Qwen3 Coder FlashQwenContext1MInputOutputProviders
+30
Throughput
Latency
Release date2025-09-17
Qwen3 Max ThinkingQwenContext262.1KInput$1.20/MOutput$6.00/MProviders
+6
48.7±16.0P
54.8±11.1E
51.7±16.1P
Throughput
Latency
Release date2026-02-09
  • 48.7
  • 54.8
  • 51.7
Qwen3.6 35B A3BQwenContext262.1KInput$0.375/MOutput$2.25/MProviders
+8
46±5.9
41.7±6.0
52.5±8.7
39.1±12.2E
34.6±11.5E
43.9±7.0
50.7±16.0P
Throughput
Latency
Release date2026-04-27
  • 46
  • 41.7
  • 52.5
  • 39.1
  • 34.6
  • 43.9
  • 50.7
Qwen3.6 Max PreviewQwenContext262.1KInput$1.30/MOutput$7.80/MProviders
+16
53.3±11.8E
50.3±8.9
54.4±10.8E
55.1±16.3P
50.2±16.1P
55±16.0P
Throughput
Latency
Release date2026-04-27
  • 53.3
  • 50.3
  • 54.4
  • 55.1
  • 50.2
  • 55
Qwen3.6 27BQwenContext262.1KInput$0.600/MOutput$3.60/MProviders
+6
55.1±6.6
47.3±6.0
54.4±8.7
37.3±11.6E
41.3±11.5E
46.1±6.5
51.8±16.0P
Throughput
Latency
Release date2026-04-27
  • 55.1
  • 47.3
  • 54.4
  • 37.3
  • 41.3
  • 46.1
  • 51.8
Qwen3.5-9BQwenContext262.1KInput$0.170/MOutput$0.250/MProviders
50.4±16.7P
51.6±13.9P
Throughput
Latency
Release date2026-03-10
  • 50.4
  • 51.6
Qwen3.5-35B-A3BQwenContext262.1KInput$0.250/MOutput$2.00/MProviders
+1
42.3±8.7
39.9±11.2E
51.1±8.3
38±16.3P
41.6±16.4P
45.8±12.1E
51.5±11.9E
Throughput
Latency
Release date2026-02-25
  • 42.3
  • 39.9
  • 51.1
  • 38
  • 41.6
  • 45.8
  • 51.5
Qwen3.5-27BQwenContext262.1KInput$0.300/MOutput$2.40/MProviders
+2
45.4±8.7
46.9±11.2E
53.7±8.3
41.4±16.3P
45.4±16.4P
49±12.1E
57.2±11.9E
Throughput
Latency
Release date2026-02-25
  • 45.4
  • 46.9
  • 53.7
  • 41.4
  • 45.4
  • 49
  • 57.2
Qwen3.5-122B-A10BQwenContext262.1KInput$0.400/MOutput$3.20/MProviders
+2
47.3±10.6E
46±11.8E
53.2±8.3
43.7±16.3P
45.4±16.4P
47.6±9.4E
54.3±11.9E
Throughput
Latency
Release date2026-02-25
  • 47.3
  • 46
  • 53.2
  • 43.7
  • 45.4
  • 47.6
  • 54.3
Qwen3 VL 32B InstructQwenContext131.1KInput$0.160/MOutput$0.640/MProviders
+3
48.3±16.0P
48.1±11.1E
49.1±16.1P
Throughput
Latency
Release date2025-10-23
  • 48.3
  • 48.1
  • 49.1
Qwen3 VL 8B ThinkingQwenContext131.1KInputOutputProviders
+3
Throughput
Latency
Release date2025-10-14
Qwen3 VL 8B InstructQwenContext262.1KInput$0.180/MOutput$0.700/MProviders
+4
44.6±16.0P
40.5±11.1E
41.5±16.1P
Throughput
Latency
Release date2025-10-14
  • 44.6
  • 40.5
  • 41.5
Qwen3 VL 30B A3B ThinkingQwenContext262.1KInputOutputProviders
+4
Throughput
Latency
Release date2025-10-06
Qwen3 VL 30B A3B InstructQwenContext262.1KInput$0.200/MOutput$0.800/MProviders
+3
47.6±16.0P
47.2±11.1E
49.8±16.1P
Throughput
Latency
Release date2025-10-06
  • 47.6
  • 47.2
  • 49.8
Qwen3 VL 235B A22B ThinkingQwenContext131.1KInputOutputProviders
+7
Throughput
Latency
Release date2025-09-23
Qwen3 VL 235B A22B InstructQwenContext262.1KInput$0.400/MOutput$1.60/MProviders
+7
49.9±16.0P
49.9±11.1E
49.5±16.1P
Throughput
Latency
Release date2025-09-23
  • 49.9
  • 49.9
  • 49.5
Qwen3 Next 80B A3B ThinkingQwenContext262.1KInputOutputProviders
+7
Throughput
Latency
Release date2025-09-11
Qwen3 Next 80B A3B InstructQwenContext262.1KInput$0.150/MOutput$1.20/MProviders
+19
51.8±16.0P
50.3±11.1E
48.7±16.1P
Throughput
Latency
Release date2025-09-11
  • 51.8
  • 50.3
  • 48.7
Qwen3 30B A3B Thinking 2507QwenContext81.9KInputOutputProviders
+6
Throughput
Latency
Release date2025-08-28
Qwen3 Coder 30B A3B InstructQwenContext262.1KInput$0.450/MOutput$2.25/MProviders
+3
46.1±16.0P
42.6±11.1E
50.5±11.8E
Throughput
Latency
Release date2025-07-31
  • 46.1
  • 42.6
  • 50.5
Qwen3 30B A3B Instruct 2507QwenContext262.1KInputOutputProviders
+5
Throughput
Latency
Release date2025-07-29
Qwen3 235B A22B Thinking 2507QwenContext131.1KInputOutputProviders
+5
Throughput
Latency
Release date2025-07-25
Qwen3 235B A22B Instruct 2507QwenContext262.1KInput$0.230/MOutput$2.30/MProviders
52.9±13.9P
53.4±11.1E
62.9±11.8E
Throughput
Latency
Release date2025-07-21
  • 52.9
  • 53.4
  • 62.9
Qwen3 30B A3BQwenContext131.1KInput$0.200/MOutput$2.40/MProviders
+1
44.4±16.0P
43±11.1E
49.4±11.8E
Throughput
Latency
Release date2025-04-28
  • 44.4
  • 43
  • 49.4
Qwen3 8BQwenContext131.1KInput$0.180/MOutput$2.10/MProviders
+1
36.7±13.9P
42.1±11.1E
48.5±11.8E
Throughput
Latency
Release date2025-04-28
  • 36.7
  • 42.1
  • 48.5
Qwen3 14BQwenContext131.1KInput$0.350/MOutput$4.20/MProviders
+2
39.8±13.9P
44.6±11.1E
49.8±11.8E
Throughput
Latency
Release date2025-04-28
  • 39.8
  • 44.6
  • 49.8
Qwen3 32BQwenContext131.1KInput$0.160/MOutput$0.640/MProviders
+2
42.8±13.9P
46.2±11.1E
49.9±11.8E
Throughput
Latency
Release date2025-04-28
  • 42.8
  • 46.2
  • 49.9
Qwen3 235B A22BQwenContext131.1KInput$0.700/MOutput$8.40/MProviders
44.8±16.0P
45.6±11.1E
51.1±11.8E
Throughput
Latency
Release date2025-04-28
  • 44.8
  • 45.6
  • 51.1
Qwen2.5 VL 72B InstructQwenContext128KInputOutputProviders
+1
Throughput
Latency
Release date2025-02-01
Qwen3 ASR FlashQwenContext0InputOutputProviders
Throughput
Latency
Release date2026-05-14
Qwen3 Embedding 8BQwenContext32.8KInputOutputProviders
+6
Throughput
Latency
Release date2025-10-28
Qwen3 Embedding 4BQwenContext32.8KInputOutputProviders
Throughput
Latency
Release date2025-10-28
Qwen-Audio-3.0-TTS PlusQwenContext0InputOutputProviders
Throughput
Latency
Release date2026-07-23
Qwen3.8 2.4T A95BQwenContext1.0MInput$2.00/MOutput$6.00/MProviders
59.2±16.0P
62.1±13.9P
Throughput
Latency
Release date2026-08-12
  • 59.2
  • 62.1
Qwen3 Reranker 8BQwenContext41.0KInputOutputProviders
+4
Throughput
Latency
Release date2026-08-13
Qwen3.8 27BQwenContext1MInput$0.500/MOutput$3.00/MProviders
+16
58.7±8.4
51.6±8.6
54.7±10.8E
43.9±14.0P
57.6±11.3E
56.3±16.0P
Throughput
Latency
Release date2026-08-14
  • 58.7
  • 51.6
  • 54.7
  • 43.9
  • 57.6
  • 56.3
Qwen3 InstantAlibabaContextInputOutputProviders
Throughput
Latency
Release date
Qwen3.5 Plus ImageContextInputOutputProviders
+3
Throughput
Latency
Release date
Qwen3.5 Plus Image EditContextInputOutputProviders
+3
Throughput
Latency
Release date
Qwen1.8B Long ContextAlibabaContextInputOutputProviders
+14
Throughput
Latency
Release date
Qwen1.8BAlibabaContextInputOutputProviders
+14
Throughput
Latency
Release date
Qwen3.5 Omni FlashContextInput$0.100/MOutput$0.800/MProviders
+22
47.5±13.9P
Throughput
Latency
Release date
  • 47.5
Qwen3.6 Plus ThinkingContextInputOutputProviders
+10
Throughput
Latency
Release date
Qwen3.5 Plus ThinkingContextInputOutputProviders
+9
Throughput
Latency
Release date
Qwen3 TTS FlashContextInputOutputProviders
+29
Throughput
Latency
Release date
Qwen3 TTS Flash RealtimeContextInputOutputProviders
+14
Throughput
Latency
Release date
Qwen3 S2S Flash RealtimeContextInputOutputProviders
+13
Throughput
Latency
Release date
Qwen ImageContextInputOutputProviders
+36
Throughput
Latency
Release date
Qwen3 NextContextInputOutputProviders
+7
Throughput
Latency
Release date
Qwen3 EmbeddingContextInputOutputProviders
+43
Throughput
Latency
Release date
FreeOpenrouterContext200KInputOutputProviders
+16
Throughput
24 t/s
Latency
25.60s
Release date2026-02-01
Qwen3.8 MaxQwenContext1MInput$2.00/MOutput$6.00/MProviders
+65
61.5±10.9E
61.7±11.2E
63.5±10.1E
52.2±14.0P
64.3±6.3
57.9±16.0P
Throughput
292 t/s
Latency
4.31s
Release date2026-08-03
  • 61.5
  • 61.7
  • 63.5
  • 52.2
  • 64.3
  • 57.9
Qwen3.8 FlashQwenContext1MInputOutputProviders
+26
Throughput
126 t/s
Latency
2.96s
Release date2026-08-26
Qwen3.7 MaxQwenContext1MInput$2.50/MOutput$7.50/MProviders
+71
56.3±5.7
55.6±6.0
59.5±8.7
55.8±12.2E
62.2±12.2E
58.1±11.7E
56.9±11.9E
Throughput
111 t/s
Latency
22.62s
Release date2026-05-21
  • 56.3
  • 55.6
  • 59.5
  • 55.8
  • 62.2
  • 58.1
  • 56.9
Qwen3.6 PlusQwenContext1MInput$0.500/MOutput$3.00/MProviders
+128
48.6±5.5
50±8.2
56.6±8.1
52±11.6E
52±9.1E
52.2±12.2E
50.2±6.4
55.7±11.9E
Throughput
49 t/s
Latency
24.59s
Release date2026-04-02
  • 48.6
  • 50
  • 56.6
  • 52
  • 52
  • 52.2
  • 50.2
  • 55.7
Qwen3.6 FlashQwenContext1MInputOutputProviders
+24
Throughput
106 t/s
Latency
8.21s
Release date2026-04-27
Qwen3.7 PlusQwenContext1MInput$0.400/MOutput$1.60/MProviders
+61
53±5.7
52.2±6.0
57.3±8.7
48.5±12.2E
55.2±12.2E
51.1±11.7E
54.7±6.3
56.9±11.9E
Throughput
80 t/s
Latency
28.78s
Release date2026-06-03
  • 53
  • 52.2
  • 57.3
  • 48.5
  • 55.2
  • 51.1
  • 54.7
  • 56.9
DeepSeek V3.2DeepSeekContext163.8KInput$0.280/MOutput$0.420/MProviders
+181
38.8±6.9
49.5±8.6
50.7±8.7
39.6±16.0P
48.7±16.1P
46.1±16.0P
Throughput
46 t/s
Latency
5.33s
Release date2025-12-01
  • 38.8
  • 49.5
  • 50.7
  • 39.6
  • 48.7
  • 46.1
Qwen3.7 FlashQwenContext1MInputOutputProviders
+18
Throughput
388 t/s
Latency
11.58s
Release date2026-07-27

How to read category scores

Agents, Coding, Reasoning, and the other capability columns are 0–100 observed-capability estimates relative to the eligible model population in a dated Category Score V3 run.

They are not success rates, IQ scores, or an average across all eight categories. Read them with the 80% interval, measured dimensions, benchmark families, and evidence shown on each model page.

Read the complete scoring methodology

What this directory shows

Each row brings together the information you need to compare an LLM API. Some fields are blank when LMSpeed has no current data for that model or provider.

API price
Compare input and output price per million tokens when it is available.
Speed and latency
Use throughput and first-token latency to compare response behavior.
Provider coverage
Open a model to review its listed providers and their current details.
Capability data
Use the capability columns when current model scores are available.

How to use the model directory

Start with the decision that matters most for your workload, then compare the current rows before you test an endpoint.

  1. Find a model.Search by model name, slug, or description.
  2. Sort the key metric.Sort by price, throughput, latency, provider count, or capability data.
  3. Compare providers.Open a model page to review the provider options and current data.
  4. Test your endpoint.Run a speed test before you use an endpoint in production.

Frequently Asked Questions

How does LMSpeed benchmark AI models?

LMSpeed runs standardized five-round API speed tests on each model, measuring output throughput (tokens per second), first-token latency, and total response time across multiple providers.

Which AI model has the lowest API latency?

Latency varies by provider and model. Use the LMSpeed model directory to sort by latency and find the model with the fastest first-token response time. Check the latency leaderboard for monthly rankings.

How to compare LLM API pricing across providers?

LMSpeed lists input and output token prices per million for each model across available providers. Sort by price to find a lower-cost option, or filter by provider to compare rates side by side.

What is an AI model directory?

An AI model directory is a searchable list of models and their comparison data. On LMSpeed, it brings together API price, throughput, first-token latency, provider coverage, and capability data when available.

How do I choose a model and provider?

Start with the metric that matters most for your workload. Then open a model page to compare provider data. Run your own speed test before you use an endpoint in production.

Can model price and speed change?

Yes. Price, availability, throughput, and latency can change by model, provider, and time. Use current page values as a comparison signal and confirm with your own endpoint test.