Sponsored byFusecodeEnterprise coding API for Claude Code, Codex, and model workflows.
LogoLMSpeed
  • Models
  • Providers
  • Free
  • Leaderboard
LogoLMSpeed
  1. Home
  2. Speed Test
LogoLMSpeed

The best API speed test tool

GitHubGitHubTwitterX (Twitter)Email
Product
  • Features
  • Pricing
  • FAQ
Leaderboard
  • Overview
  • Speed Ranking
  • Latency Ranking
  • Health Ranking
  • Model Pricing
  • Model Speed
  • Reasoning
  • Coding
Models
  • All Models
  • GPT
  • Claude
  • Gemini
  • DeepSeek
  • Llama
  • Qwen
Free Models
  • All Free Models
  • Free GPT
  • Free Claude
  • Free Gemini
  • Free DeepSeek
  • Free Llama
  • Free Qwen
Tools
  • Speed Test
  • Provider Audit
Company
  • About
Resources
  • Provider Directory
  • Documentation
  • Public API
  • Botab
  • VidBee
Legal
  • Cookie Policy
  • Privacy Policy
  • Terms of Service
© 2026 LMSpeed All Rights Reserved.Made by Nexmoe with ❤️

LLM API benchmark

Test how fast your AI API really responds

Run the same five prompts against any compatible endpoint, then compare the delay before output starts, generation speed, and total response time.

What Speedtest measures

Measure the wait, then measure the flow

LMSpeed sends five standardized prompts to the model and records each response from the first token to completion. This separates connection and queue delay from the model's sustained generation speed.

Why this is useful

Use comparable results to spot slow relay routes, evaluate providers before integration, and choose an endpoint that feels responsive to real users.

01First-token latency
How long a user waits before the first visible token arrives. Lower is better for chat and interactive products.
02Output throughput
How many tokens the API generates each second after output begins. Higher means long answers finish sooner.
03Total duration
The complete time for each standardized response, combining startup delay and generation time.

Recent Test Results

A live sample of recent community benchmarks for comparing provider latency and output speed.

TimeProviderModelSpeedLatency
Aug 15, 03:05 AMmishradev.com
claude-sonnet-5
85.74t/s
2.74s
Aug 14, 06:31 PMopen.bigmodel.cn
glm-5.3
420.40t/s
27.86s
Aug 14, 02:55 PMintegrate.api.nvidia.com
stepfun-ai/step-3.7-flash
267.40t/s
37.55s
Aug 14, 02:55 PMintegrate.api.nvidia.com
meta/muse-glimmer-30b
118.96t/s
4.57s
Aug 14, 02:53 PMintegrate.api.nvidia.com
meta/muse-glimmer-30b
129.74t/s
5.10s

Query model scores and validate the API behind them

Start with model score lookup, then compare per-token pricing, five-round speed tests, health checks, and safety probes before an API reaches production.

Model Benchmark Lookup

Search a model name from the homepage and jump into benchmark pages with AA score, coding, math, and model metadata.

API Pricing Comparison

Compare per-token pricing across 100+ providers, find cheaper APIs for each model, and track free tiers and credits.

Real-time Speed Benchmarks

Run a five-round benchmark with standardized prompts to measure first-token latency, output throughput, and response time.

API Security Audit

Audit any OpenAI-compatible API for model authenticity, hidden prompts, instruction tampering, stream integrity, and error leakage, then share a plain-language report.

Custom Endpoint Benchmarks

Enter a base URL, API key, and model ID to test official providers, proxies, relays, or self-hosted endpoints.

Speed Benchmark Analytics

Review first-token latency, output throughput, total duration, health, and recent probe signals to judge stability.

Frequently Asked Questions

How LMSpeed handles model benchmarks, provider pricing, speed tests, and API audits.

How do I query model benchmark scores?

Use the model search on the homepage or open the Model Performance leaderboard. Search by model name to find the model detail page, where LMSpeed connects benchmark signals such as AA score, coding, and math with model metadata.

How do I compare LLM API pricing across providers?

LMSpeed aggregates per-token pricing from 100+ API providers. Visit any model page to see a side-by-side pricing comparison table showing input and output rates per million tokens, so you can find the cheapest provider for each model.

Which LLM APIs are free?

Many providers offer free API tiers or credits for popular models like DeepSeek, Gemini, and Llama. Check our Free LLM API directory for a complete list of models with free access, including speed benchmarks for each free provider.

How does LMSpeed conduct speed benchmark testing?

LMSpeed employs a five-round continuous stress testing mechanism with standardized prompts. Token calculations are performed accurately using tiktoken, measuring output throughput (tokens per second) and first-token latency.

What does the API trust audit check?

It sends multiple safety probes to an OpenAI-compatible endpoint to check whether the model identity matches, hidden system prompts are injected, user instructions are rewritten, streaming responses stay intact, and errors leak sensitive implementation details. The result includes a risk score and a shareable report. API keys are only used for that audit and are not written to public reports.

How to compare speed between different API providers?

Use our performance leaderboards and model detail pages to visually compare API speed benchmarks across providers. The system ranks providers by throughput, latency, and health, helping you choose the fastest and most reliable API.

Is long-term performance monitoring supported?

Provider pages and the health leaderboard already show recent health checks, probe latency, success or failure status, and stability rankings. Broader continuous monitoring and alerting will keep expanding.