Choose a model to open its comparison page.
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workfl...
Input and output token limits for this model, plus how it ranks on long-context understanding.
Third-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
DeepSeek deepseek/fp8 | $0.140/M | $0.280/M | 100.0% | — | — | 1.0M tokens / 384K tokens |
GMICloud gmicloud/fp8 | $0.140/M | $0.280/M | 100.0% | — | — | 1.0M tokens / — |
Cloudflare cloudflare/fp8 | $0.140/M | $0.280/M | 99.9% | — | — | 384K tokens / 384K tokens |
SiliconFlow siliconflow/fp8 | $0.140/M | $0.280/M | 99.8% | — | — | 1.0M tokens / 393.2K tokens |
DeepInfra deepinfra/fp4 | $0.090/M | $0.180/M | 95.3% | — | — | 1.0M tokens / 65.5K tokens |
Parasail parasail/fp8 | $0.140/M | $0.280/M | 92.5% | — | — | 1.0M tokens / 1.0M tokens |
AtlasCloud atlas-cloud/fp8 | $0.140/M | $0.280/M | 84.2% | — | — | 262.1K tokens / 131.1K tokens |
Fireworks fireworks | $0.140/M | $0.280/M | 66.1% | — | — | 1.0M tokens / — |