Review Agent, Coding, and Reasoning model scores, compare provider rates, free tiers, and model coverage, measure API latency, throughput, and duration, and detect model, prompt, and error leakage risks.
Start with model score lookup, then compare per-token pricing, five-round speed tests, health checks, and safety probes before an API reaches production.
Search a model name from the homepage and jump into benchmark pages with AA score, coding, math, and model metadata.
Compare per-token pricing across 100+ providers, find cheaper APIs for each model, and track free tiers and credits.
Run a five-round benchmark with standardized prompts to measure first-token latency, output throughput, and response time.
Audit any OpenAI-compatible API for model authenticity, hidden prompts, instruction tampering, stream integrity, and error leakage, then share a plain-language report.
Enter a base URL, API key, and model ID to test official providers, proxies, relays, or self-hosted endpoints.
Review first-token latency, output throughput, total duration, health, and recent probe signals to judge stability.
How LMSpeed handles model benchmarks, provider pricing, speed tests, and API audits.
Use the model search on the homepage or open the Model Performance leaderboard. Search by model name to find the model detail page, where LMSpeed connects benchmark signals such as AA score, coding, and math with model metadata.
LMSpeed aggregates per-token pricing from 100+ API providers. Visit any model page to see a side-by-side pricing comparison table showing input and output rates per million tokens, so you can find the cheapest provider for each model.
Many providers offer free API tiers or credits for popular models like DeepSeek, Gemini, and Llama. Check our Free LLM API directory for a complete list of models with free access, including speed benchmarks for each free provider.
LMSpeed employs a five-round continuous stress testing mechanism with standardized prompts. Token calculations are performed accurately using tiktoken, measuring output throughput (tokens per second) and first-token latency.
It sends multiple safety probes to an OpenAI-compatible endpoint to check whether the model identity matches, hidden system prompts are injected, user instructions are rewritten, streaming responses stay intact, and errors leak sensitive implementation details. The result includes a risk score and a shareable report. API keys are only used for that audit and are not written to public reports.
Use our performance leaderboards and model detail pages to visually compare API speed benchmarks across providers. The system ranks providers by throughput, latency, and health, helping you choose the fastest and most reliable API.
Provider pages and the health leaderboard already show recent health checks, probe latency, success or failure status, and stability rankings. Broader continuous monitoring and alerting will keep expanding.
The newest relay audit reports where endpoint profile, model identity, prompt safety, and response integrity all scored 100.
A live cut of newly tracked models and benchmark leaders, focused on Artificial Analysis scores for overall intelligence, coding, and math.
| Model | Context | Input | Output | Providers | Agents | Coding | Reasoning | Knowledge | Math | Multilingual | Multimodal | Instruction following | Throughput | Latency | Release date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5Anthropic | Context1M | Input$10.00/M | Output$50.00/M | Providers +105 | 66.4±8.2 | 67.8±6.9 | 61±10.8E | 60.8±14.0P | — | — | 45.6±17.0P | 50.6±16.0P | Throughput 58 t/s | Latency 3.77s | Release date2026-06-09 |
| Kimi K3MoonshotAINEW | Context1.0M | Input$3.00/M | Output$15.00/M | Providers +7 | 69.3±7.3 | 61.1±11.9E | 61.7±10.8E | 61.5±14.0P | — | — | 69.7±11.3E | — | Throughput — | Latency — | Release date2026-07-16 |
| GPT-5.6 SolOpenAINEW | Context1.1M | Input$5.00/M | Output$30.00/M | Providers +101 | 68±7.3 | 62.5±9.3E | 61.7±8.7 | 57.2±16.7P | 67±16.1P | — | 60.9±16.1P | 53.9±16.0P | Throughput 53 t/s | Latency 2.06s | Release date2026-07-09 |
| Claude Opus 4.8Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +138 | 63.4±5.3 | 65.7±6.6 | 56.7±8.7 | 63.9±14.0P | 58±16.1P | 55±17.1P | 62.7±12.0E | 50.2±16.0P | Throughput 232 t/s | Latency 2.21s | Release date2026-05-27 |
| GPT-5.5OpenAI | Context1.1M | Input$5.00/M | Output$30.00/M | Providers +123 | 62.3±5.1 | 60.1±8.7 | 59.6±8.7 | 60±14.0P | 56.9±16.1P | — | 56.4±16.1P | 55.1±16.0P | Throughput 45 t/s | Latency 5.42s | Release date2026-04-24 |
| Grok 4.5xAINEW | Context500K | Input$2.00/M | Output$6.00/M | Providers +60 | 62.4±8.4 | 59±6.9 | 54.5±8.7 | 57.9±14.0P | — | — | 48.5±17.0P | — | Throughput 54 t/s | Latency 5.37s | Release date2026-07-08 |
| Claude Opus 4.7Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +185 | 55.1±6.9 | 60.5±11.2E | 56.5±10.8E | 53.7±14.0P | 56.8±16.1P | — | — | 44.5±16.0P | Throughput 47 t/s | Latency 4.89s | Release date2026-05-12 |
| Claude Sonnet 5AnthropicNEW | Context1M | Input$2.00/M | Output$10.00/M | Providers +87 | 60.7±8.2 | 59.3±6.6 | 56.6±10.8E | 61.3±14.0P | — | — | 62.3±16.2P | — | Throughput — | Latency — | Release date2026-06-30 |
| GPT-5.6 LunaOpenAINEW | Context1.1M | Input$1.00/M | Output$6.00/M | Providers +92 | 60.4±7.4 | 57.2±9.3E | 54.9±8.7 | 56.5±16.7P | 62.4±16.1P | — | 50.1±16.1P | — | Throughput — | Latency — | Release date2026-07-09 |
| GLM-5.2Z.ai | Context1.0M | Input$1.40/M | Output$4.40/M | Providers +122 | 60.9±7.2 | 58.1±8.9 | 54.9±10.8E | 58.3±14.0P | 70.4±11.5E | — | — | 54.1±16.0P | Throughput 61 t/s | Latency 7.79s | Release date2026-06-16 |
| Gemini 3.5 FlashGoogle | Context1.0M | Input$1.50/M | Output$9.00/M | Providers +89 | 61.3±5.6 | 56.1±8.9 | 53.4±8.4 | 55.2±14.0P | 55.1±16.1P | — | 58.5±11.9E | 55.3±16.0P | Throughput 425 t/s | Latency 4.10s | Release date2026-05-19 |
| Claude Sonnet 4.6Anthropic | Context1M | Input$3.00/M | Output$15.00/M | Providers +212 | 53.7±6.1 | 56.2±8.6 | 50.2±8.7 | 70.6±16.3P | 53.2±16.1P | — | 44.5±16.2P | 43.8±16.0P | Throughput 43 t/s | Latency 4.13s | Release date2026-02-17 |
| Gemini 3.1 Pro PreviewGoogle | Context1.0M | Input$2.00/M | Output$12.00/M | Providers +8 | — | 66.5±16.0P | 64.6±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-02-19 |
| Qwen3.7 MaxQwen | Context1M | Input$2.50/M | Output$7.50/M | Providers +36 | 57.2±5.4 | 59.4±6.0 | 59.9±8.7 | 55.8±12.2E | 64.5±12.4E | 58.1±11.7E | — | 57.1±11.9E | Throughput 68 t/s | Latency 17.93s | Release date2026-05-21 |
| MiniMax M3MiniMax | Context1.0M | Input$0.300/M | Output$1.20/M | Providers +68 | 56.3±7.3 | 53.5±6.6 | 59.4±10.8E | 52.6±14.0P | — | — | 47.1±12.0E | 58.4±16.0P | Throughput 65 t/s | Latency 2.03s | Release date2026-05-31 |
| GPT-5.3 CodexOpenAI | Context400K | Input$1.75/M | Output$14.00/M | Providers +243 | 53.3±6.6 | 58.9±8.5 | 60.5±10.8E | 55±16.0P | — | — | — | 54.9±16.0P | Throughput 78 t/s | Latency 3.91s | Release date2026-02-24 |
Compare API pricing, speed benchmarks, and performance data across providers.
Advertising
The first provider slot is open for sponsorship.
A unified API gateway providing access to multiple large language models with direct connectivity in China.
Health
100%
Tests
140
Last check
Jul 29
API price
No health checks yet
Health
98%
Tests
10
Last check
Jul 29
API price
No health checks yet
Health
100%
Tests
45
Last check
Jul 29
API price
No health checks yet
Health
100%
Tests
15
Last check
Jul 29
API price
No health checks yet
xAI provides the Grok series of AI models through its API, offering text generation and multimodal capabilities.
Health
69%
Tests
45
Last check
Jul 29
API price
No health checks yet
NVIDIA NIM provides optimized AI model inference APIs for LLMs, vision, and embedding models through NVIDIA cloud infrastructure.
Health
100%
Tests
1,184
Last check
Jul 29
API price
No health checks yet
DeepSeek provides API access to its latest large language models for text generation and coding tasks.
Health
100%
Tests
622
Last check
Jul 29
API price
No health checks yet
Hanbing API is a community welfare relay with daily check-in rewards, suited for LLM and tavern chat. High-consumption tools like OpenClaw and Codex are not allowed.
Health
100%
Tests
40
Last check
Jul 29
API price
No health checks yet
CM-API (api.chengmo.cc.cd) is a LinuxDO LLM API relay by user chengmo. 0.01 USD per call. Grok, Kimi, Qwen. Supports immersive translate and LDC.
Health
100%
Tests
5
Last check
Jul 29
API price
No health checks yet
iFlytek Spark MaaS platform offering Spark series LLMs with strong Chinese language capabilities via OpenAI-compatible API.
Health
100%
Tests
108
Last check
Jul 29
API price
No health checks yet
ChooseC API is a unified AI model aggregation gateway supporting 260+ mainstream models including Claude, GPT, Qwen, DeepSeek, Kimi, and GLM with OpenAI, Claude, and Gemini compatibility.
Health
100%
Tests
110
Last check
Jul 29
API price
No health checks yet
OpenCode is an open-source AI coding agent that integrates with terminals, IDEs, and desktop apps, supporting multiple models and providers.
Health
100%
Tests
140
Last check
Jul 29
API price
No health checks yet