DeepSeek Chat V3.1 API Benchmarks, Pricing & Provider Data
Compare DeepSeek Chat V3.1 with another model
Choose a model to open its comparison page.
The page also shows measured API speed and first-token latency.
DeepSeek Chat V3.1 is an instruction-tuned variant in the DeepSeek series, optimized for following instructions and conversational tasks.
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Released
- Aug 2025
- Knowledge cutoff
- 2025-03-31
- Tokenizer
- DeepSeek
- Architecture
- text->text
- Instruct type
- deepseek-v3.1
- Moderated
- No
- Supported parameters
- frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
OpenRouter endpoints
8 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
DeepInfra deepinfra/fp4 | $0.250/M | $0.950/M | 99.9% | — | — | undefined tokens / undefined tokens |
Novita novita/fp8 | $0.270/M | $1/M | 99.9% | — | — | undefined tokens / undefined tokens |
CoreWeave coreweave/fp8 | $0.550/M | $1.65/M | 99.9% | — | — | undefined tokens / undefined tokens |
AtlasCloud atlas-cloud/fp8 | $0.300/M | $0.950/M | 99.6% | — | — | undefined tokens / undefined tokens |
SambaNova sambanova/fp8 | $0.650/M | $1.50/M | 98.4% | — | — | undefined tokens / undefined tokens |
SiliconFlow siliconflow/fp8 | $0.270/M | $1/M | 97.0% | — | — | undefined tokens / undefined tokens |
Mara mara | $0.600/M | $1.70/M | 95.9% | — | — | undefined tokens / undefined tokens |
Google google-vertex/us-west2 | $0.600/M | $1.70/M | 0% | — | — | undefined tokens / undefined tokens |
Alternatives & Similar Models
Gemini 2.5 Pro
gemini-2-5-pro
Google Gemini 2.5 Pro is Google advanced multimodal model with a 1M-token context window, strong STEM reasoning, and native support for images, audio, and video understanding.
Qwen3
qwen3
Alibaba Qwen3 is the Qwen family's flagship LLM series with dense and MoE variants, seamless thinking/non-thinking modes, and leading open-source performance in math, code, and agent tasks.
Kimi K2
kimi-k2
Moonshot Kimi K2 is a trillion-parameter MoE model from Moonshot AI, built for long-context chat, agentic tool use, and high-quality Chinese-English bilingual generation.
Qwen3 Coder
qwen3-coder
Alibaba Qwen3 Coder is a code-specialized variant in the Qwen series, optimized for code generation, debugging, and software development tasks.
GPT-5.4
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
GPT-5.3 Codex
gpt-5-3-codex
OpenAI GPT-5.3 Codex is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks.
