Query model scores and test AI APIs before you choose
Review Agent, Coding, and Reasoning model scores, compare provider rates, free tiers, and model coverage, measure API latency, throughput, and duration, and detect model, prompt, and error leakage risks.
Our sponsors
Partners supporting independent model and API provider benchmarks.
Latest 100/100 API Security Audits
The newest relay audit reports where endpoint profile, model identity, prompt safety, and response integrity all scored 100.
- sub.callai.onegpt-5.6-solReport timeAug 18Report timeAug 18848486100
- sub.callai.onegpt-5.6-solReport timeAug 18Report timeAug 18707280100
- ai.databyte.co.iddatabyte-m1Report timeAug 18Report timeAug 187284100100
- ai.databyte.co.iddeepseek-v4-flashReport timeAug 18Report timeAug 1810084100100
- ai.databyte.co.idMiniMax-M3Report timeAug 18Report timeAug 18768478100
- tokengate-cqt9ivzs.manus.spaceclaude-opus-5Report timeAug 17Report timeAug 17648465100
- router.bynara.idqwen-3.8-max-freeReport timeAug 17Report timeAug 17846686100
- 87.120.187.161claude-fable-5Report timeAug 17Report timeAug 1772727273
- muyuan.doGLM-5.2Report timeAug 17Report timeAug 17768468100
- code28.ccwu.ccGPT-5.6 TerraReport timeAug 17Report timeAug 17848486100
Latest and strongest LLM models
A live cut of newly tracked models and benchmark leaders, focused on Artificial Analysis scores for overall intelligence, coding, and math.
| Model | Context | Input | Output | Providers | Agents | Coding | Reasoning | Knowledge | Math | Multilingual | Multimodal | Instruction following | Throughput | Latency | Release date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1AnthropicNEW | Context1M | Input$10.00/M | Output$50.00/M | Providers +65 | 61.7±8.6 | 68.5±6.9 | 61.2±8.7 | 68±14.0P | — | — | — | — | Throughput — | Latency — | Release date2026-09-01 |
| Claude Opus 5Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +106 | 65.9±7.7 | 67.5±6.6 | 59.5±8.7 | 66.9±14.0P | — | 64.3±17.1P | 56.7±16.4P | — | Throughput 284 t/s | Latency 5.44s | Release date2026-07-24 |
| Claude Fable 5Anthropic | Context1M | Input$10.00/M | Output$50.00/M | Providers +132 | 63.1±8.2 | 66.5±6.9 | 60.9±10.8E | 64.9±14.0P | — | — | 44.8±16.4P | 50.4±16.0P | Throughput 61 t/s | Latency 3.92s | Release date2026-06-09 |
| GPT-6 AstraOpenAINEW | Context1.1M | Input$10.00/M | Output$50.00/M | Providers +87 | 63±8.4 | 59.5±11.9E | 62.1±8.7 | 58±16.6P | 72.3±16.1P | — | 66.4±16.4P | — | Throughput — | Latency — | Release date2026-09-04 |
| Muse Spark 1.3MetaNEW | Context1.0M | Input$1.25/M | Output$4.25/M | Providers +2 | 63.4±8.4 | 58.3±11.9E | 60.9±10.8E | 56.9±14.0P | — | — | — | — | Throughput — | Latency — | Release date2026-09-02 |
| GLM 5.3Z.aiNEW | Context1.3M | Input$1.40/M | Output$4.40/M | Providers +80 | 50.8±18.4P | 62.9±16.0P | 60.8±13.9P | — | — | — | — | — | Throughput 379 t/s | Latency 30.00s | Release date2026-08-18 |
| Grok 4.6SpaceXAI | Context500K | Input$2.00/M | Output$6.00/M | Providers +92 | 60.8±8.4 | 58.5±11.9E | 59.5±10.8E | 60.2±14.0P | — | — | — | — | Throughput 150 t/s | Latency 22.76s | Release date2026-08-12 |
| Kimi K3MoonshotAI | Context1.0M | Input$3.00/M | Output$15.00/M | Providers +115 | 65.7±7.3 | 56.4±11.9E | 60.1±10.8E | 59.4±14.0P | — | — | 64.9±11.3E | — | Throughput 138 t/s | Latency 30.06s | Release date2026-07-16 |
| GPT-5.6 SolOpenAI | Context1.1M | Input$4.00/M | Output$20.00/M | Providers +152 | 65±7.3 | 60.8±9.3E | 60.9±8.7 | 55.9±16.6P | 66±16.1P | — | 59.5±16.1P | 53.5±16.0P | Throughput 53 t/s | Latency 4.44s | Release date2026-07-09 |
| Claude Opus 4.8Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +163 | 61.8±5.5 | 64±6.6 | 55.9±8.7 | 62.1±14.0P | 57.2±16.1P | 55±17.1P | 59±12.0E | 50±16.0P | Throughput 193 t/s | Latency 4.02s | Release date2026-05-27 |
| GLM 5.3 FlashZ.aiNEW | Context1.3M | Input$0.150/M | Output$0.500/M | Providers +78 | 60.6±11.9E | 56.6±9.3E | 57.6±11.1E | 54.2±14.0P | — | — | 57.9±16.1P | — | Throughput 336 t/s | Latency 31.71s | Release date2026-08-26 |
| Gemini 3.8 FlashGoogleNEW | Context1.0M | Input$0.750/M | Output$3.75/M | Providers +60 | 59.9±11.3E | 58.8±11.9E | 59.7±10.8E | 61.3±14.0P | — | — | 54±16.4P | — | Throughput — | Latency — | Release date2026-09-02 |
| Claude Opus 4.7Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +202 | 53.7±6.9 | 58.5±11.2E | 54.8±10.8E | 53.6±14.0P | 56±16.1P | — | — | 44.5±16.0P | Throughput 47 t/s | Latency 4.89s | Release date2026-05-12 |
| Qwen3.8 MaxQwen | Context1M | Input$2.00/M | Output$6.00/M | Providers +64 | 61.5±10.9E | 61.7±11.2E | 63.5±10.1E | 52.2±14.0P | — | — | 64.3±6.3 | 57.9±16.0P | Throughput 292 t/s | Latency 4.31s | Release date2026-08-03 |
| Qwen3.8 2.4T A95BQwen | Context1.0M | Input$2.00/M | Output$6.00/M | Providers | — | 59.2±16.0P | 62.1±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-12 |
| Muse Spark 1.2Meta | Context1.0M | Input$1.25/M | Output$4.25/M | Providers +2 | 56.3±8.9 | 55.7±11.9E | 58.5±10.8E | 58.8±14.0P | — | — | — | — | Throughput 69 t/s | Latency 22.53s | Release date2026-08-05 |
Query model scores and validate the API behind them
Start with model score lookup, then compare per-token pricing, five-round speed tests, health checks, and safety probes before an API reaches production.
- Model Benchmark Lookup
Search a model name from the homepage and jump into benchmark pages with AA score, coding, math, and model metadata.
- API Pricing Comparison
Compare per-token pricing across 100+ providers, find cheaper APIs for each model, and track free tiers and credits.
- Real-time Speed Benchmarks
Run a five-round benchmark with standardized prompts to measure first-token latency, output throughput, and response time.
- API Security Audit
Audit any OpenAI-compatible API for model authenticity, hidden prompts, instruction tampering, stream integrity, and error leakage, then share a plain-language report.
- Custom Endpoint Benchmarks
Enter a base URL, API key, and model ID to test official providers, proxies, relays, or self-hosted endpoints.
- Speed Benchmark Analytics
Review first-token latency, output throughput, total duration, health, and recent probe signals to judge stability.
Frequently Asked Questions
How LMSpeed handles model benchmarks, provider pricing, speed tests, and API audits.
How do I query model benchmark scores?
Use the model search on the homepage or open the Model Performance leaderboard. Search by model name to find the model detail page, where LMSpeed connects benchmark signals such as AA score, coding, and math with model metadata.
How do I compare LLM API pricing across providers?
LMSpeed aggregates per-token pricing from 100+ API providers. Visit any model page to see a side-by-side pricing comparison table showing input and output rates per million tokens, so you can find the cheapest provider for each model.
Which LLM APIs are free?
Many providers offer free API tiers or credits for popular models like DeepSeek, Gemini, and Llama. Check our Free LLM API directory for a complete list of models with free access, including speed benchmarks for each free provider.
How does LMSpeed conduct speed benchmark testing?
LMSpeed employs a five-round continuous stress testing mechanism with standardized prompts. Token calculations are performed accurately using tiktoken, measuring output throughput (tokens per second) and first-token latency.
What does the API trust audit check?
It sends multiple safety probes to an OpenAI-compatible endpoint to check whether the model identity matches, hidden system prompts are injected, user instructions are rewritten, streaming responses stay intact, and errors leak sensitive implementation details. The result includes a risk score and a shareable report. API keys are only used for that audit and are not written to public reports.
How to compare speed between different API providers?
Use our performance leaderboards and model detail pages to visually compare API speed benchmarks across providers. The system ranks providers by throughput, latency, and health, helping you choose the fastest and most reliable API.
Is long-term performance monitoring supported?
Provider pages and the health leaderboard already show recent health checks, probe latency, success or failure status, and stability rankings. Broader continuous monitoring and alerting will keep expanding.
