LMSpeed is an AI model directory for comparing API price, output speed, first-token latency, provider coverage, and benchmark data. Use it to narrow your model and provider choice, then test the endpoint that fits your workload.
Price, speed, latency, and availability can change. Treat this table as a current signal and verify your own endpoint before you deploy.
| Release date | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Context1.0M | Input$0.750/M |
Agents, Coding, Reasoning, and the other capability columns are 0–100 observed-capability estimates relative to the eligible model population in a dated Category Score V3 run.
They are not success rates, IQ scores, or an average across all eight categories. Read them with the 80% interval, measured dimensions, benchmark families, and evidence shown on each model page.
Each row brings together the information you need to compare an LLM API. Some fields are blank when LMSpeed has no current data for that model or provider.
Start with the decision that matters most for your workload, then compare the current rows before you test an endpoint.
LMSpeed runs standardized five-round API speed tests on each model, measuring output throughput (tokens per second), first-token latency, and total response time across multiple providers.
Latency varies by provider and model. Use the LMSpeed model directory to sort by latency and find the model with the fastest first-token response time. Check the latency leaderboard for monthly rankings.
LMSpeed lists input and output token prices per million for each model across available providers. Sort by price to find a lower-cost option, or filter by provider to compare rates side by side.
| Providers |
63.6±14.1P |
61.7±16.0P |
59.5±13.9P |
— |
— |
— |
59.5±16.1P |
— |
| Throughput — | Latency — |
| Release date2026-08-13 |
|
| Qwen3 Reranker 8BQwen | Context41.0K | Input— | Output— | Providers | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-13 |
|
| Qwen3.8 2.4T A95BQwen | Context1.0M | Input$2.00/M | Output$6.00/M | Providers— | — | 60.4±16.0P | 62.7±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-12 |
|
| Grok 4.6SpaceXAI | Context500K | Input$2.00/M | Output$6.00/M | Providers +4 | 62.3±8.6 | 60.5±12.0E | 59.9±10.8E | 59.8±16.0P | — | — | — | — | Throughput 97 t/s | Latency 10.46s | Release date2026-08-12 |
|
| DeepSeek V4 Pro 0813DeepSeek | Context1.0M | Input— | Output— | Providers— | 57.4±7.7 | 57.1±6.1 | 59.5±8.3 | 58.4±18.0P | 58.9±12.2E | — | — | 55±16.0P | Throughput — | Latency — | Release date2026-08-12 |
|
| Nemotron 3.5 LightningNVIDIA | Context1M | Input$0.050/M | Output$0.200/M | Providers | — | 46.4±16.0P | 49.1±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-11 |
|
| Solar Pro 4Upstage | Context524.3K | Input$0.300/M | Output$1.20/M | Providers | — | 55±16.0P | 58.1±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-10 |
|
| Muse Glimmer 30BMeta | Context131.1K | Input— | Output— | Providers— | 48.3±10.2E | 48.6±9.3E | 56.3±10.8E | 43.9±16.0P | 46.1±16.3P | — | 46.3±9.4E | 55.2±16.0P | Throughput — | Latency — | Release date2026-08-09 |
|
| Ling 3.0 TinyinclusionAI | Context262.1K | Input— | Output— | Providers | — | 40.2±16.0P | 48.4±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-06 |
|
| Muse Spark 1.2Meta | Context1.0M | Input$1.25/M | Output$4.25/M | Providers | 56.9±9.4E | 57.2±12.0E | 61.5±10.8E | 55.8±14.0P | — | — | — | — | Throughput — | Latency — | Release date2026-08-05 |
|
| Qwen3.8 MaxQwen | Context1M | Input$2.00/M | Output$6.00/M | Providers +14 | 66.5±10.9E | 66.1±11.2E | 64.7±10.1E | 52.7±14.0P | — | — | 65.5±6.3 | 57.8±16.0P | Throughput 93 t/s | Latency 3.28s | Release date2026-08-03 |
|
| Inkling SmallThinking Machines | Context524.3K | Input$0.300/M | Output$1.20/M | Providers | 55.8±10.9E | 52.7±6.9 | 51±8.7 | 52.5±14.0P | 52.5±14.4P | — | 46±11.9E | 57.5±16.0P | Throughput — | Latency — | Release date2026-07-30 |
|
| DeepSeek V4 Flash 0731DeepSeek | Context1.0M | Input— | Output— | Providers | 52.1±8.4 | 52.6±6.3 | 59.5±8.3 | 47.5±18.0P | 57.9±12.2E | — | — | — | Throughput 74 t/s | Latency 2.70s | Release date2026-07-31 |
|
| H3MiniMax | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-29 |
|
| Gen-4.5Runway | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-29 |
|
| Aleph 2.0Runway | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-29 |
|
| S2.1 ProFish Audio | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-29 |
|
| S2.1 Pro FreeFish Audio | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-29 |
|
| S2 ProFish Audio | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-29 |
|
| Transcribe 1Fish Audio | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-29 |
|
| voyage-multimodal-3.5VoyageAI by MongoDB | Context32K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-27 |
|
| rerank-2.5VoyageAI by MongoDB | Context32K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-27 |
|
| rerank-2.5-liteVoyageAI by MongoDB | Context32K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-27 |
|
| MiniMax M3 (batch)MiniMax | Context524.3K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-05-31 |
|
| Gemini 2.5 Pro (batch)Google | Context1.0M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-06-17 |
|
| Gemini 2.5 Flash (batch)Google | Context1.0M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-06-17 |
|
| Gemini 2.5 Flash Lite (batch)Google | Context1.0M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-07-22 |
|
| Gemini 3 Flash Preview (batch)Google | Context1.0M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-12-17 |
|
| Gemini 3.1 Pro Preview (batch)Google | Context1.0M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-02-19 |
|
| Gemini 3.1 Flash Lite (batch)Google | Context1.0M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-05-07 |
|
| Gemini 3.5 Flash (batch)Google | Context1.0M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-05-19 |
|
| Gemini 3.5 Flash Lite (batch)Google | Context1.0M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-21 |
|
| Gemini 3.6 Flash (batch)Google | Context1.0M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-21 |
|
| Claude Opus 5 (batch)Anthropic | Context1M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-24 |
|
| Claude Opus 4.1 (batch)Anthropic | Context200K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-08-05 |
|
| Claude Sonnet 4.5 (batch)Anthropic | Context1M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-09-29 |
|
| Claude Haiku 4.5 (batch)Anthropic | Context200K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-10-15 |
|
| Text Embedding 3 Small (batch)OpenAI | Context8.2K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-10-30 |
|
| Text Embedding 3 Large (batch)OpenAI | Context8.2K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-10-30 |
|
| Text Embedding Ada 002 (batch)OpenAI | Context8.2K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-10-30 |
|
| Claude Opus 4.5 (batch)Anthropic | Context200K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-11-24 |
|
| Claude Opus 4.6 (batch)Anthropic | Context1M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-02-04 |
|
| Claude Opus 4.7 (batch)Anthropic | Context1M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-04-16 |
|
| Claude Opus 4.8 (batch)Anthropic | Context1M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-05-27 |
|
| Claude Fable 5 (batch)Anthropic | Context1M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-06-09 |
|
| Claude Sonnet 5 (batch)Anthropic | Context1M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-06-30 |
|
| Qwen3.7 FlashQwen | Context1M | Input— | Output— | Providers +2 | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-27 |
|
| Claude Opus 5Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +79 | 65.9±8.2 | 68.4±6.6 | 60.5±8.7 | 67.4±14.0P | — | 64.3±17.1P | 57.6±17.3P | — | Throughput 320 t/s | Latency 4.96s | Release date2026-07-24 |
|
| Claude Opus 5 (Fast)Anthropic | Context1M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-24 |
|
| Grok STT 1.0SpaceXAI | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-23 |
|
| Ling-3.0-flashInclusionai | Context262.1K | Input$0.075/M | Output$0.220/M | Providers +8 | 48.2±10.2E | 50.2±9.3E | 52.8±10.8E | 38.2±14.0P | 46.7±11.5E | — | — | 54.2±16.0P | Throughput — | Latency — | Release date2026-07-23 |
|
| MAI-Voice-2-FlashMicrosoft | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-23 |
|
| MAI-Image-2.5 ProMicrosoft | Context4.1K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-23 |
|
| Qwen-Audio-3.0-TTS PlusQwen | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-23 |
|
| Qwen-Audio-3.0-TTS FlashQwen | Context0 | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2026-07-23 |
|
| GPT-3.5 Turbo (batch)OpenAI | Context16.4K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2023-05-28 |
|
| GPT-4 Turbo (batch)OpenAI | Context128K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-04-09 |
|
| GPT-4o (batch)OpenAI | Context128K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-05-13 |
|
| GPT-4o-mini (batch)OpenAI | Context128K | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2024-07-18 |
|
| GPT-4.1 Nano (batch)OpenAI | Context1.0M | Input— | Output— | Providers— | — | — | — | — | — | — | — | — | Throughput — | Latency — | Release date2025-04-14 |
|
An AI model directory is a searchable list of models and their comparison data. On LMSpeed, it brings together API price, throughput, first-token latency, provider coverage, and capability data when available.
Start with the metric that matters most for your workload. Then open a model page to compare provider data. Run your own speed test before you use an endpoint in production.
Yes. Price, availability, throughput, and latency can change by model, provider, and time. Use current page values as a comparison signal and confirm with your own endpoint test.