AI Model Benchmark Directory

LMSpeed is an AI model directory for comparing API price, output speed, first-token latency, provider coverage, and benchmark data. Use it to narrow your model and provider choice, then test the endpoint that fits your workload.

Price, speed, latency, and availability can change. Treat this table as a current signal and verify your own endpoint before you deploy.

ModelContextInputOutputProvidersAgentsCodingReasoningKnowledgeMathMultilingualMultimodalInstruction followingThroughputLatencyRelease date
Gemini 3.7 FlashGoogleContext1.0MInput$0.750/MOutput$3.75/MProviders
+7
57.2±8.6
58.5±11.9E
59.4±10.8E
57.3±14.0P
57.2±16.1P
Throughput
Latency
Release date2026-08-13
  • 57.2
  • 58.5
  • 59.4
  • 57.3
  • 57.2
Gemini 3.5 Flash-LiteGoogleContext1.0MInput$0.300/MOutput$2.50/MProviders
+38
43.6±8.8
47.2±9.3E
53.1±10.8E
51.4±14.0P
45.8±17.3P
Throughput
Latency
Release date2026-07-21
  • 43.6
  • 47.2
  • 53.1
  • 51.4
  • 45.8
Gemini 3.6 FlashGoogleContext1.0MInput$0.750/MOutput$3.75/MProviders
+65
53.3±8.6
55.7±11.9E
59.7±10.8E
57.3±14.0P
50.6±17.3P
Throughput
737 t/s
Latency
2.34s
Release date2026-07-21
  • 53.3
  • 55.7
  • 59.7
  • 57.3
  • 50.6
Gemini Embedding 2 PreviewGoogleContext8.2KInputOutputProviders
Throughput
Latency
Release date2026-04-17
Gemini 3.1 Flash TTS PreviewGoogleContext32.8KInputOutputProviders
+7
Throughput
Latency
Release date2026-04-24
Gemini 2.5 Pro Preview 05-06GoogleContext1.0MInputOutputProviders
Throughput
Latency
Release date2025-05-07
Gemini 2.5 Pro Preview 06-05GoogleContext1.0MInput$1.25/MOutput$10.00/MProviders
54±13.9P
55.4±11.1E
60.7±11.8E
Throughput
Latency
Release date2025-06-05
  • 54
  • 55.4
  • 60.7
Gemini 2.5 Flash Lite Preview 09-2025GoogleContext1.0MInput$0.100/MOutput$0.400/MProviders
46.3±13.9P
47.8±11.1E
45.1±16.1P
Throughput
Latency
Release date
  • 46.3
  • 47.8
  • 45.1
Nano Banana Pro (Gemini 3 Pro Image Preview)GoogleContext65.5KInputOutputProviders
+3
Throughput
Latency
Release date2025-11-20
Gemini 3 Flash PreviewGoogleContext1.0MInput$0.500/MOutput$3.00/MProviders
+5
62.4±13.9P
60.2±11.1E
54.4±16.1P
Throughput
Latency
Release date2025-12-17
  • 62.4
  • 60.2
  • 54.4
Gemini 3.1 Pro PreviewGoogleContext1.0MInput$2.00/MOutput$12.00/MProviders
+11
65.2±16.0P
63.8±13.9P
Throughput
Latency
Release date2026-02-19
  • 65.2
  • 63.8
Gemini 3.1 Pro Preview Custom ToolsGoogleContext1.0MInputOutputProviders
+9
Throughput
Latency
Release date2026-02-25
Nano Banana 2 (Gemini 3.1 Flash Image Preview)GoogleContext65.5KInputOutputProviders
+4
Throughput
Latency
Release date2026-02-26
Gemini 3.1 Flash Lite PreviewGoogleContext1.0MInput$0.250/MOutput$1.50/MProviders
+5
53.8±16.0P
53.1±13.9P
Throughput
Latency
Release date2026-03-03
  • 53.8
  • 53.1
Lyria 3 Clip PreviewGoogleContext1.0MInputOutputProviders
Throughput
Latency
Release date2026-03-30
Lyria 3 Pro PreviewGoogleContext1.0MInputOutputProviders
Throughput
Latency
Release date2026-03-30
Gemini 3.1 Flash Lite ImageGoogleContext65.5KInputOutputProviders
+32
Throughput
Latency
Release date2026-06-30
Gemini 3.5 FlashGoogleContext1.0MInput$1.50/MOutput$9.00/MProviders
+107
60.6±5.6
55.9±8.9
53.4±8.4
54.5±14.0P
55.1±16.1P
56.3±11.9E
54.8±16.0P
Throughput
425 t/s
Latency
4.10s
Release date2026-05-19
  • 60.6
  • 55.9
  • 53.4
  • 54.5
  • 55.1
  • 56.3
  • 54.8
Gemini 2.5 Pro DeepSearchGoogleContextInputOutputProviders
+2
Throughput
Latency
Release date
Gemini 3.0 Pro ImageGoogleContextInputOutputProviders
Throughput
Latency
Release date
Gemini 1.5 Pro 002GoogleContextInputOutputProviders
+2
Throughput
Latency
Release date
Gemini 2.5 Flash DeepSearchGoogleContextInputOutputProviders
+2
Throughput
Latency
Release date
Gemini 3.0 ProGoogleContextInputOutputProviders
+1
Throughput
Latency
Release date
Gemini Robotics ER 1.5GoogleContextInputOutputProviders
+19
Throughput
Latency
Release date
Gemini 3 Pro DeepSearchGoogleContextInputOutputProviders
+1
Throughput
Latency
Release date
Gemini 2.5 Flash LiveGoogleContextInputOutputProviders
Throughput
Latency
Release date
Gemini 2.5 Pro 1MGoogleContextInputOutputProviders
Throughput
Latency
Release date
Gemini 3.1 FlashGoogleContextInputOutputProviders
+19
Throughput
Latency
Release date
Gemini EmbeddingGoogleContextInputOutputProviders
+8
Throughput
Latency
Release date
Gemini 2.0 Flash Live 001GoogleContextInputOutputProviders
+1
Throughput
Latency
Release date
Gemini Pro VisionGoogleContextInputOutputProviders
Throughput
Latency
Release date
Gemini Live 2.5 FlashGoogleContextInputOutputProviders
Throughput
Latency
Release date
Gemini 3.1 ProGoogleContext1.0MInputOutputProviders
+188
50.6±6.1
58.8±10.7E
58.3±8.7
57.8±15.1P
55.3±16.1P
59.7±17.3P
55.4±6.6
55.1±16.0P
Throughput
92 t/s
Latency
15.36s
Release date2026-02-19
  • 50.6
  • 58.8
  • 58.3
  • 57.8
  • 55.3
  • 59.7
  • 55.4
  • 55.1
Gemini Embedding 2GoogleContext8.2KInputOutputProviders
+36
Throughput
Latency
Release date2026-05-20
Gemini 1.0 Pro VisionGoogleContextInputOutputProviders
Throughput
Latency
Release date
Gemini 1.5 Flash 002GoogleContextInputOutputProviders
+1
Throughput
155 t/s
Latency
1.62s
Release date
Gemini 3.0 FlashGoogleContextInputOutputProviders
+4
Throughput
28 t/s
Latency
20.33s
Release date
Gemini 1.5 FlashGoogleContextInputOutputProviders
+18
33.3±13.9P
37.2±11.1E
42.6±11.8E
Throughput
200 t/s
Latency
1.17s
Release date
  • 33.3
  • 37.2
  • 42.6
Gemini ProGoogleContext1.0MInputOutputProviders
+41
Throughput
Latency
Release date2026-04-27
Gemini 2.0 ProGoogleContextInputOutputProviders
+12
43.8±13.9P
48.2±11.1E
52.1±11.8E
Throughput
67 t/s
Latency
8.47s
Release date
  • 43.8
  • 48.2
  • 52.1
Gemini 2.0 Flash LiteGoogleContextInputOutputProviders
+40
37.3±13.9P
41.5±13.9P
50.1±11.8E
Throughput
182 t/s
Latency
1.49s
Release date
  • 37.3
  • 41.5
  • 50.1
Gemini 2.0 Flash Lite 001GoogleContextInputOutputProviders
+22
37.6±13.9P
43.3±11.1E
49.8±11.8E
Throughput
Latency
Release date
  • 37.6
  • 43.3
  • 49.8
Gemini 2.0 FlashGoogleContextInputOutputProviders
+64
44.7±13.9P
46.5±11.1E
52.1±11.8E
Throughput
137 t/s
Latency
3.08s
Release date
  • 44.7
  • 46.5
  • 52.1
Gemini Embedding 001GoogleContext20KInputOutputProviders
+59
Throughput
Latency
Release date2025-10-31
Gemini FlashGoogleContext1.0MInputOutputProviders
+45
Throughput
Latency
Release date2026-04-27
Gemini 3 Pro ImageGoogleContext131.1KInputOutputProviders
+112
Throughput
Latency
Release date2026-06-18
Gemini 2.0 Flash 001GoogleContextInputOutputProviders
+24
Throughput
108 t/s
Latency
2.70s
Release date
Gemini 2.5 Flash Native AudioGoogleContextInputOutputProviders
+7
Throughput
Latency
Release date
Gemini 3.1 Flash ImageGoogleContext131.1KInputOutputProviders
+133
Throughput
Latency
Release date2026-06-18
Gemini 2.5 FlashGoogleContext1.0MInput$0.300/MOutput$2.50/MProviders
+170
36.3±16.0P
50.6±13.9P
49±8.7
40.9±16.0P
50.8±9.3E
43.1±16.0P
Throughput
155 t/s
Latency
8.55s
Release date2025-06-17
  • 36.3
  • 50.6
  • 49
  • 40.9
  • 50.8
  • 43.1
Gemini 3 FlashGoogleContext1.0MInput$0.500/MOutput$3.00/MProviders
+185
43±8.9
53.7±11.2E
52.2±8.7
50.6±16.0P
51.9±16.1P
47.8±16.0P
Throughput
162 t/s
Latency
7.48s
Release date2025-12-17
  • 43
  • 53.7
  • 52.2
  • 50.6
  • 51.9
  • 47.8
Gemini Flash LiteGoogleContextInputOutputProviders
+45
Throughput
Latency
Release date
Gemini 1.5 Pro 001GoogleContextInputOutputProviders
+1
Throughput
Latency
Release date
Gemini 3.1 Flash LiteGoogleContext1.0MInputOutputProviders
+128
44.5±16.1P
27.5±16.1P
42.6±16.1P
Throughput
191 t/s
Latency
6.58s
Release date2026-05-07
  • 44.5
  • 27.5
  • 42.6
Gemini 2.5 Flash LiteGoogleContext1.0MInput$0.100/MOutput$0.400/MProviders
+128
36.3±13.9P
42.8±11.1E
53.3±11.8E
Throughput
277 t/s
Latency
1.41s
Release date2025-09-25
  • 36.3
  • 42.8
  • 53.3
Gemini 2.5 Flash ImageGoogleContext32.8KInputOutputProviders
+93
Throughput
Latency
Release date2025-10-07
Gemini 1.5 ProGoogleContextInputOutputProviders
+21
40.1±13.9P
39.6±11.1E
54.2±16.0P
43.7±11.8E
Throughput
18 t/s
Latency
2.58s
Release date
  • 40.1
  • 39.6
  • 54.2
  • 43.7
Gemini 3 ProGoogleContextInput$2.00/MOutput$12.00/MProviders
+139
49±9.3E
55.6±11.2E
54.1±6.4
56.1±14.0P
55.7±16.1P
54±17.3P
49.9±6.5
52.6±16.0P
Throughput
73 t/s
Latency
10.59s
Release date
  • 49
  • 55.6
  • 54.1
  • 56.1
  • 55.7
  • 54
  • 49.9
  • 52.6
Gemini 2.5 ProGoogleContext1.0MInput$1.25/MOutput$10.00/MProviders
+171
43.6±9.3E
39.1±8.9
54.2±8.7
35±14.0P
56.1±9.3E
45.9±16.0P
Throughput
88 t/s
Latency
16.94s
Release date2025-06-17
  • 43.6
  • 39.1
  • 54.2
  • 35
  • 56.1
  • 45.9

How to read category scores

Agents, Coding, Reasoning, and the other capability columns are 0–100 observed-capability estimates relative to the eligible model population in a dated Category Score V3 run.

They are not success rates, IQ scores, or an average across all eight categories. Read them with the 80% interval, measured dimensions, benchmark families, and evidence shown on each model page.

Read the complete scoring methodology

What this directory shows

Each row brings together the information you need to compare an LLM API. Some fields are blank when LMSpeed has no current data for that model or provider.

API price
Compare input and output price per million tokens when it is available.
Speed and latency
Use throughput and first-token latency to compare response behavior.
Provider coverage
Open a model to review its listed providers and their current details.
Capability data
Use the capability columns when current model scores are available.

How to use the model directory

Start with the decision that matters most for your workload, then compare the current rows before you test an endpoint.

  1. Find a model.Search by model name, slug, or description.
  2. Sort the key metric.Sort by price, throughput, latency, provider count, or capability data.
  3. Compare providers.Open a model page to review the provider options and current data.
  4. Test your endpoint.Run a speed test before you use an endpoint in production.

Frequently Asked Questions

How does LMSpeed benchmark AI models?

LMSpeed runs standardized five-round API speed tests on each model, measuring output throughput (tokens per second), first-token latency, and total response time across multiple providers.

Which AI model has the lowest API latency?

Latency varies by provider and model. Use the LMSpeed model directory to sort by latency and find the model with the fastest first-token response time. Check the latency leaderboard for monthly rankings.

How to compare LLM API pricing across providers?

LMSpeed lists input and output token prices per million for each model across available providers. Sort by price to find a lower-cost option, or filter by provider to compare rates side by side.

What is an AI model directory?

An AI model directory is a searchable list of models and their comparison data. On LMSpeed, it brings together API price, throughput, first-token latency, provider coverage, and capability data when available.

How do I choose a model and provider?

Start with the metric that matters most for your workload. Then open a model page to compare provider data. Run your own speed test before you use an endpoint in production.

Can model price and speed change?

Yes. Price, availability, throughput, and latency can change by model, provider, and time. Use current page values as a comparison signal and confirm with your own endpoint test.