AI Model Benchmark Directory
LMSpeed is an AI model directory for comparing API price, output speed, first-token latency, provider coverage, and benchmark data. Use it to narrow your model and provider choice, then test the endpoint that fits your workload.
Price, speed, latency, and availability can change. Treat this table as a current signal and verify your own endpoint before you deploy.
| Model | Context | Input | Output | Providers | Agents | Coding | Reasoning | Knowledge | Math | Multilingual | Multimodal | Instruction following | Throughput | Latency | Release date | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V3 | Context— | Input$0.360/M | Output$0.890/M | Providers +151 | 35.2±11.8E | 37.3±11.2E | 44.3±8.7 | 40.9±16.0P | 49.5±9.3E | — | — | 39.7±11.9E | Throughput 49 t/s | Latency 3.39s | Release date— |
|
| Qwen3 | Context262.1K | Input$0.200/M | Output$0.800/M | Providers +126 | — | 45.7±13.9P | 47.6±11.1E | 36.9±16.3P | 58.3±11.8E | 36.7±16.4P | — | — | Throughput 100 t/s | Latency 12.55s | Release date2025-07-21 |
|
| Qwen3 VL | Context— | Input— | Output— | Providers +17 | — | — | — | — | — | — | — | — | Throughput 206 t/s | Latency 6.63s | Release date— |
|
| Devstral 2 | Context— | Input— | Output— | Providers +4 | — | 46.1±13.9P | 45.2±11.1E | — | 43.2±16.1P | — | — | — | Throughput 71 t/s | Latency 1.34s | Release date— |
|
| Ministral 3 | Context— | Input$0.200/M | Output$0.200/M | Providers +2 | — | 31.8±13.9P | 36.5±11.1E | — | 40.5±16.1P | — | — | — | Throughput 208 t/s | Latency 2.84s | Release date— |
|
| Mistral Large 3 | Context— | Input$0.500/M | Output$1.50/M | Providers +7 | 39.6±11.8E | 47.8±13.9P | 44.5±8.7 | 41.5±16.0P | 43.5±16.1P | — | — | 42.2±16.0P | Throughput 29 t/s | Latency 3.00s | Release date— |
|
| Qwen3 CoderQwen | Context262.1K | Input— | Output— | Providers +61 | — | — | — | — | — | — | — | — | Throughput 243 t/s | Latency 5.28s | Release date2025-07-23 |
|
| Qwen3 Embedding | Context— | Input— | Output— | Providers +43 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| Qwen3 Next | Context— | Input— | Output— | Providers +7 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| DeepSeek Coder | Context— | Input— | Output— | Providers +8 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| Qwen Image | Context— | Input— | Output— | Providers +34 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| Phi 4Microsoft | Context16.4K | Input$0.125/M | Output$0.500/M | Providers +17 | 28.3±16.0P | 39.3±13.9P | 39.1±8.7 | 37.4±16.0P | 46.9±11.8E | — | — | 37.8±16.0P | Throughput 43 t/s | Latency 1.32s | Release date2025-01-10 |
|
| Intern-S1 | Context— | Input— | Output— | Providers +4 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| MiniMax M1MiniMax | Context1M | Input$0.550/M | Output$2.20/M | Providers +31 | — | 51.1±13.9P | 49.3±11.1E | — | 58.9±11.8E | — | — | — | Throughput — | Latency — | Release date2025-06-17 |
|
| MAI-DS-R1 | Context— | Input— | Output— | Providers +10 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| DeepSeek VL2 | Context— | Input— | Output— | Providers +8 | — | — | — | — | — | — | — | — | Throughput 126 t/s | Latency 0.75s | Release date— |
|
| GLM-4 | Context— | Input— | Output— | Providers +57 | — | — | — | — | — | — | — | — | Throughput 61 t/s | Latency 0.87s | Release date— |
|
| GPT-OSS | Context131.1K | Input— | Output— | Providers +137 | — | — | — | — | — | — | — | — | Throughput 374 t/s | Latency 3.23s | Release date2025-08-05 |
|
| GPT-4oOpenAI | Context128K | Input$2.50/M | Output$10.00/M | Providers +133 | 39±16.0P | 46±13.9P | 39.9±8.7 | 48.9±16.0P | 42.2±9.3E | — | — | 41.6±16.0P | Throughput 82 t/s | Latency 3.44s | Release date2024-11-20 |
|
| MiniMax M2.5MiniMax | Context204.8K | Input$0.300/M | Output$1.20/M | Providers +188 | — | 50±11.8E | 54.5±13.9P | — | — | — | — | — | Throughput 61 t/s | Latency 9.68s | Release date2026-02-12 |
|
| GPT-5.4 MiniOpenAI | Context400K | Input$0.750/M | Output$4.50/M | Providers +240 | 50.9±8.1 | 57.2±11.8E | 55.5±10.8E | 46.1±14.0P | 49.7±16.1P | — | 47.1±16.1P | 53.6±16.0P | Throughput 134 t/s | Latency 3.54s | Release date2026-03-17 |
|
| MiniMax Hailuo 2.3 | Context— | Input— | Output— | Providers +16 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date— |
|
| GLM-5Z.ai | Context204.8K | Input$1.00/M | Output$3.20/M | Providers +188 | 41±5.6 | 51.1±8.1 | 48.8±8.1 | 43.3±16.3P | 51.1±9.1E | 43.8±12.2E | — | 52.4±11.9E | Throughput 51 t/s | Latency 21.59s | Release date2026-02-11 |
|
| GPT-4o MiniOpenAI | Context128K | Input$0.150/M | Output$0.600/M | Providers +121 | 32±16.0P | 37.5±13.9P | 40.2±11.1E | 53.2±16.0P | 46.1±11.8E | — | — | 40.5±16.0P | Throughput 84 t/s | Latency 4.10s | Release date2024-07-18 |
|
| Claude Sonnet 4.5Anthropic | Context1M | Input$3.00/M | Output$15.00/M | Providers +174 | 46.5±8.7 | 51.6±11.2E | 45.5±8.9E | 54.2±16.1P | 42.5±12.3E | — | — | — | Throughput 41 t/s | Latency 4.57s | Release date2025-09-29 |
|
| GPT-5.1OpenAI | Context400K | Input$1.25/M | Output$10.00/M | Providers +153 | 47.7±9.3E | 49.1±11.2E | 53.9±8.7 | 52.9±16.0P | 54±16.1P | — | — | 53.5±16.0P | Throughput 142 t/s | Latency 2.78s | Release date2025-11-13 |
|
| GLM-4.6VZ.ai | Context131.1K | Input$0.300/M | Output$0.900/M | Providers +61 | — | 40.1±13.9P | 49.7±11.1E | — | 52.2±16.1P | — | — | — | Throughput 40 t/s | Latency 27.37s | Release date2025-12-08 |
|
| Gemini 2.5 ProGoogle | Context1.0M | Input$1.25/M | Output$10.00/M | Providers +172 | 43.8±9.3E | 38.9±8.9 | 54.1±8.7 | 36.3±14.0P | 56.1±9.3E | — | — | 45.9±16.0P | Throughput 88 t/s | Latency 16.94s | Release date2025-06-17 |
|
| O1OpenAI | Context200K | Input$15.00/M | Output$60.00/M | Providers +75 | 45.5±16.0P | 50.5±13.9P | 50.7±8.7 | 62.5±17.4P | 52.6±9.3E | — | — | 51.5±11.9E | Throughput — | Latency — | Release date2024-12-17 |
|
| O1 MiniOpenAI | Context— | Input— | Output— | Providers +60 | — | 47.4±13.9P | 44.6±11.1E | — | 54.9±11.8E | — | — | — | Throughput 32 t/s | Latency 13.84s | Release date— |
|
| GPT-4.1OpenAI | Context1.0M | Input$2.00/M | Output$8.00/M | Providers +109 | 40±11.8E | 38.1±11.2E | 48.5±8.7 | 57±17.4P | 49±9.3E | — | — | 42.1±11.9E | Throughput 85 t/s | Latency 1.92s | Release date2025-04-14 |
|
| O3OpenAI | Context200K | Input$2.00/M | Output$8.00/M | Providers +85 | 49.3±16.0P | 55.1±13.9P | 54.3±8.7 | 47.6±16.0P | 58.8±9.3E | — | — | 53±16.0P | Throughput 134 t/s | Latency 2.82s | Release date2025-04-16 |
|
| Kimi K2MoonshotAI | Context131.1K | Input$0.570/M | Output$2.30/M | Providers +87 | 45.3±16.0P | 48.2±13.9P | 48.5±8.7 | 44.3±16.0P | 53.7±9.3E | — | — | 43.8±16.0P | Throughput 29 t/s | Latency 2.21s | Release date2025-09-04 |
|
| Claude Opus 4.6Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +229 | 54.8±5.7 | 58.6±8.0 | 53.3±8.7 | 53.7±13.7P | 56.5±16.1P | 54±17.3P | 48.4±11.1E | 44.7±16.0P | Throughput 45 t/s | Latency 5.30s | Release date2026-02-04 |
|
| Gemini 3 ProGoogle | Context— | Input$2.00/M | Output$12.00/M | Providers +139 | 49±9.3E | 55.3±11.2E | 53.9±6.4 | 55.1±14.0P | 55.7±16.1P | 54±17.3P | 49.9±6.5 | 52.6±16.0P | Throughput 73 t/s | Latency 10.59s | Release date— |
|
| Phi 4 Mini Instruct | Context131.1K | Input— | Output— | Providers +28 | — | 27.8±13.9P | 34.3±11.1E | — | 42±11.8E | — | — | — | Throughput — | Latency — | Release date2025-10-17 |
|
| GPT-5 MiniOpenAI | Context400K | Input$0.250/M | Output$2.00/M | Providers +111 | — | 49.4±11.2E | 53.5±11.1E | — | 52.2±16.1P | — | — | — | Throughput 90 t/s | Latency 7.03s | Release date2025-08-07 |
|
| GPT-5.1 Codex MaxOpenAI | Context400K | Input— | Output— | Providers +126 | 49.9±16.0P | 50.7±11.8E | 54.2±10.8E | 49.9±16.0P | — | — | — | 52.5±16.0P | Throughput 56 t/s | Latency 6.29s | Release date2025-12-04 |
|
| Gemini 1.5 ProGoogle | Context— | Input— | Output— | Providers +21 | — | 40.3±13.9P | 39.6±11.1E | 54±16.0P | 43.7±11.8E | — | — | — | Throughput 18 t/s | Latency 2.58s | Release date— |
|
| GPT-4.1 MiniOpenAI | Context1.0M | Input$0.400/M | Output$1.60/M | Providers +112 | 40.7±11.8E | 38.8±11.2E | 45.3±8.7 | 49.3±17.4P | 49.9±9.3E | — | — | 42.4±11.9E | Throughput 93 t/s | Latency 2.69s | Release date2025-04-14 |
|
| Grok Imagine VideoSpaceXAI | Context0 | Input— | Output— | Providers +22 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-05-18 |
|
| Claude Opus 4Anthropic | Context200K | Input$15.00/M | Output$75.00/M | Providers +91 | — | 50.9±13.9P | 51.6±11.1E | — | 54.4±11.8E | — | — | — | Throughput 25 t/s | Latency 12.05s | Release date2025-05-22 |
|
| Gemini 2.5 Flash ImageGoogle | Context32.8K | Input— | Output— | Providers +93 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-10-07 |
|
| GPT-5.4OpenAI | Context1.1M | Input$2.50/M | Output$15.00/M | Providers +294 | 57±5.1 | 62.1±10.3E | 55.9±8.7 | 59.9±15.1P | 57.6±16.1P | — | 52.9±6.6 | 53.9±16.0P | Throughput 49 t/s | Latency 4.45s | Release date2026-03-05 |
|
| GPT Audio MiniOpenAI | Context128K | Input— | Output— | Providers +16 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-01-19 |
|
| GLM-4.7Z.ai | Context204.8K | Input$0.600/M | Output$2.20/M | Providers +152 | 46.6±8.6 | 52.8±10.5E | 54.4±8.7 | 34.1±14.0P | 48.8±12.3E | — | — | 51.8±16.0P | Throughput 95 t/s | Latency 19.09s | Release date2025-12-22 |
|
| Claude Opus 4.5Anthropic | Context200K | Input$5.00/M | Output$25.00/M | Providers +166 | 48.4±5.3 | 55.2±8.1 | 60.6±8.1 | 57.6±11.6E | 49.9±9.1E | 51.7±12.2E | 32.4±6.5 | 45.7±11.9E | Throughput 54 t/s | Latency 3.12s | Release date2025-11-24 |
|
| GLM-4.6Z.ai | Context204.8K | Input$0.550/M | Output$2.20/M | Providers +102 | 48.3±16.0P | 41.9±11.2E | 42.8±8.7 | 43.5±16.0P | 44.8±16.1P | — | — | 42.4±16.0P | Throughput 53 t/s | Latency 16.68s | Release date2025-09-30 |
|
| Gemini 2.5 Flash LiteGoogle | Context1.0M | Input$0.100/M | Output$0.400/M | Providers +128 | — | 36.7±13.9P | 42.8±11.1E | — | 53.3±11.8E | — | — | — | Throughput 277 t/s | Latency 1.41s | Release date2025-09-25 |
|
| DeepSeek R1DeepSeek | Context64K | Input$1.35/M | Output$3.00/M | Providers +149 | 41.2±16.0P | 49.6±13.9P | 50.1±8.7 | 44.6±16.0P | 57±11.8E | — | — | 43.3±16.0P | Throughput 52 t/s | Latency 10.36s | Release date2025-05-28 |
|
| Devstral SmallMistral AI | Context— | Input— | Output— | Providers +7 | — | 38.9±13.9P | 39.4±11.1E | — | 43.4±11.8E | — | — | — | Throughput — | Latency — | Release date— |
|
| GPT-3.5 Turbo InstructOpenAI | Context4.1K | Input— | Output— | Providers +44 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2023-09-28 |
|
| GPT-3.5 TurboOpenAI | Context16.4K | Input$0.500/M | Output$1.50/M | Providers +98 | — | 48.8±18.4P | 33.3±11.8E | — | 37.8±16.0P | — | — | — | Throughput 115 t/s | Latency 2.21s | Release date2024-01-25 |
|
| Gemini 3.1 Flash LiteGoogle | Context1.0M | Input— | Output— | Providers +128 | 44.5±16.1P | 27.4±16.1P | — | — | — | — | 42.6±16.1P | — | Throughput 191 t/s | Latency 6.58s | Release date2026-05-07 |
|
| O3 ProOpenAI | Context200K | Input$20.00/M | Output$80.00/M | Providers +61 | — | — | 54±16.0P | 60±16.0P | — | — | — | — | Throughput — | Latency — | Release date2025-06-10 |
|
| GPT-4 TurboOpenAI | Context128K | Input$10.00/M | Output$30.00/M | Providers +73 | — | 43.3±13.9P | 43.2±11.8E | 53.5±16.0P | 45.8±11.8E | — | — | — | Throughput 50 t/s | Latency 0.97s | Release date2024-04-09 |
|
| GPT-5.1 Codex MiniOpenAI | Context400K | Input$0.250/M | Output$2.00/M | Providers +125 | — | 56.4±13.9P | 53.1±11.1E | — | 53.4±16.1P | — | — | — | Throughput 136 t/s | Latency 4.28s | Release date2025-11-13 |
|
| Claude Opus 4.1Anthropic | Context200K | Input$15.00/M | Output$75.00/M | Providers +100 | 46.4±16.2P | 46.9±16.1P | — | 58.9±16.0P | — | — | — | — | Throughput 22 t/s | Latency 3.56s | Release date2025-08-05 |
|
| Qwen3 Omni Flash | Context— | Input— | Output— | Providers +39 | — | — | — | — | — | — | — | — | Throughput 62 t/s | Latency 5.13s | Release date— |
|
| DeepSeek ReasonerDeepSeek | Context— | Input— | Output— | Providers +85 | — | — | — | — | — | — | — | — | Throughput 28 t/s | Latency 6.40s | Release date— |
|
How to read category scores
Agents, Coding, Reasoning, and the other capability columns are 0–100 observed-capability estimates relative to the eligible model population in a dated Category Score V3 run.
They are not success rates, IQ scores, or an average across all eight categories. Read them with the 80% interval, measured dimensions, benchmark families, and evidence shown on each model page.
What this directory shows
Each row brings together the information you need to compare an LLM API. Some fields are blank when LMSpeed has no current data for that model or provider.
- API price
- Compare input and output price per million tokens when it is available.
- Speed and latency
- Use throughput and first-token latency to compare response behavior.
- Provider coverage
- Open a model to review its listed providers and their current details.
- Capability data
- Use the capability columns when current model scores are available.
How to use the model directory
Start with the decision that matters most for your workload, then compare the current rows before you test an endpoint.
- Find a model.Search by model name, slug, or description.
- Sort the key metric.Sort by price, throughput, latency, provider count, or capability data.
- Compare providers.Open a model page to review the provider options and current data.
- Test your endpoint.Run a speed test before you use an endpoint in production.
Frequently Asked Questions
How does LMSpeed benchmark AI models?
LMSpeed runs standardized five-round API speed tests on each model, measuring output throughput (tokens per second), first-token latency, and total response time across multiple providers.
Which AI model has the lowest API latency?
Latency varies by provider and model. Use the LMSpeed model directory to sort by latency and find the model with the fastest first-token response time. Check the latency leaderboard for monthly rankings.
How to compare LLM API pricing across providers?
LMSpeed lists input and output token prices per million for each model across available providers. Sort by price to find a lower-cost option, or filter by provider to compare rates side by side.
What is an AI model directory?
An AI model directory is a searchable list of models and their comparison data. On LMSpeed, it brings together API price, throughput, first-token latency, provider coverage, and capability data when available.
How do I choose a model and provider?
Start with the metric that matters most for your workload. Then open a model page to compare provider data. Run your own speed test before you use an endpoint in production.
Can model price and speed change?
Yes. Price, availability, throughput, and latency can change by model, provider, and time. Use current page values as a comparison signal and confirm with your own endpoint test.
