Choose a model to open its comparison page.
DeepSeek V4 Flash benchmark, API pricing, and provider data cover 381 API providers, with prices starting at $0.0000014/M. DeepSeek V4 Flash free API options are available from 9 providers. The page also shows measured API speed and first-token latency.
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reason...
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
Input and output token limits for this model, plus how it ranks on long-context understanding.
2 metrics
2 metrics
17 metrics · Rated
15 metrics · Rated
7 metrics · Rated
9 metrics · Provisional
5 metrics · Estimated
0 metrics · No data
1 metric · No data
0 metrics · No data
Third-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
Cloudflare cloudflare | $0.140/M | $0.280/M | 100.0% | — | — | 384K tokens / 384K tokens |
DeepInfra deepinfra/fp4 | $0.090/M | $0.180/M | 99.7% | — | — | 1.0M tokens / 65.5K tokens |
Baidu baidu/fp8 | $0.069/M | $0.137/M | 99.6% | — | — | 1.0M tokens / 131.1K tokens |
DeepSeek deepseek | $0.140/M | $0.280/M | 99.6% | — | — | 1.0M tokens / 384K tokens |
DigitalOcean digitalocean | $0.068/M | $0.168/M | 99.4% | — | — | 1.0M tokens / — |
GMICloud gmicloud/fp8 | $0.094/M | $0.188/M | 99.4% | — | — | 1.0M tokens / — |
Venice venice | $0.138/M | $0.275/M | 99.4% | — | — | 1M tokens / 32.8K tokens |
Novita novita/fp8 | $0.140/M | $0.280/M | 99.3% | — | — | 1.0M tokens / 393.2K tokens |
AtlasCloud atlas-cloud/fp4 | $0.140/M | $0.280/M | 99.0% | — | — | 1.0M tokens / 393.2K tokens |
StreamLake streamlake/fp8 | $0.068/M | $0.137/M | 98.9% | — | — | 1.0M tokens / 384K tokens |
CoreWeave coreweave/fp8 | $0.140/M | $0.280/M | 98.7% | — | — | 1.0M tokens / 1.0M tokens |
Alibaba alibaba/fp8 | $0.134/M | $0.268/M | 98.7% | — | — | 1M tokens / 393.2K tokens |
Parasail parasail/fp8 | $0.140/M | $0.280/M | 98.3% | — | — | 1.0M tokens / 1.0M tokens |
SiliconFlow siliconflow/fp8 | $0.130/M | $0.280/M | 97.9% | — | — | 1.0M tokens / 393.2K tokens |
Fireworks fireworks | $0.140/M | $0.280/M | 97.7% | — | — | 1.0M tokens / — |
Morph morph | $0.139/M | $0.278/M | 95.8% | — | — | 1.0M tokens / 1.0M tokens |
Mancer 2 mancer/fp4 | $0.175/M | $0.500/M | 92.4% | — | — | 1.0M tokens / 1.0M tokens |
OpenInference open-inference/fp8 | $0.070/M | $0.180/M | 92.0% | — | — | 1.0M tokens / 393.2K tokens |
Phala phala | $0.200/M | $0.400/M | 91.8% | — | — | 1.0M tokens / 393.2K tokens |
Ambient ambient/fp4 | $0.140/M | $0.280/M | 0% | — | — | 1.0M tokens / 1.0M tokens |
Compare DeepSeek V4 Flash API pricing across 372 providers. Prices range from $0.0000014/M to $1398.60/M. AIGCBAR offers the lowest rate at $0.0000014/M. 9 providers offer free API credits or a free tier.
deepseek-v4-pro
DeepSeek V4 Pro is the professional-tier DeepSeek V4 model, targeting frontier reasoning, coding, and agent workflows with maximum capability.
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
glm-5-1
Zhipu GLM-5.1 is a next-generation GLM model aimed at frontier reasoning, coding, and bilingual agent applications.
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
deepseek-v3-2
DeepSeek V3.2 is an upgraded V3-series MoE model with stronger reasoning, coding, and math performance, widely available through OpenAI-compatible API relays.
claude-sonnet-4-6
Anthropic Claude Sonnet 4.6 extends the Sonnet line with improved tool use, coding reliability, and long-context performance for everyday production workloads.