Qwen3.8 Flash API Benchmarks, Pricing & Provider Data
Compare Qwen3.8 Flash with another model
Choose a model to open its comparison page.
Qwen3.8 Flash API pricing covers undefined API provider} other undefined API providers}}, from $0.000015/M to $100.00/request. Qwen3.8 Flash free API options are available from undefined provider} other undefined providers}}.
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart anal...
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Released
- Aug 2026
- Tokenizer
- Qwen
- Architecture
- text+image+video->text
- Moderated
- No
- Supported parameters
- frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Pricing Comparison
Compare Qwen3.8 Flash API pricing across 34 providers. Prices range from $0.000015/M to $100.00/request. OAI2API offers the lowest rate at $0.000015/M. 2 providers offer free API credits or a free tier.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% | qwen3.8-flash | bailian | $0.057/M Cache read$0.0071/M | $0.193/M | — | — | — | |
L1 100% | qwen3.8-flash | qwen | $0.320/M | $1.08/M | — | — | — | |
L1 100% L2 100% | qwen3.8-flash | Token | $0.150/M Cache read$0.016/M | $0.470/M | — | — | — | |
L1 100% L2 0% | qwen3.8-flash | default | $13.00/request | - | — | — | — | |
L1 100% | qwen3.8-flash | qwen-new | $0.099/M Cache read$0.012/MCache write$0.154/MCache write 1h$0.247/M | $0.333/M | — | — | — | |
L1 100% | qwen3.8-flash | Qwen | $0.137/M Cache read$0.014/M | $0.411/M | — | — | — | |
L1 100% L2 100% | qwen3.8-flash | default | $0.800/M Cache read$0.100/M | $2.70/M | — | — | — | |
L1 99% | qwen3.8-flash | gpt | $75.00/M | $75.00/M | — | — | — | |
L1 100% L2 100% | qwen3.8-flash | default | $37.50/M | $37.50/M | — | — | — | |
L1 99% L2 100% | qwen3.8-flash | diamond-glm | $54.75/M | $54.75/M | — | — | — | |
L1 99% | qwen3.8-flash | free | $0.000015/M Cache write$0.0000016/MCache write 1h$0.00000256/M | $0.000047/M | — | — | — | |
L1 100% | qwen3.8-flash | Alibaba-2 | $0.033/M Cache read$0.0036/M | $0.105/M | — | — | — | |
L1 100% | qwen3.8-flash | default | $0.500/M Cache read$0.050/MCache write$0.625/MCache write 1h$1.00/M | $1.50/M | — | — | — | |
L1 99% | qwen3.8-flash | 国产 | $30.00/M | $30.00/M | — | — | — | |
L1 100% | qwen3.8-flash | Alibaba-2 | $0.033/M Cache read$0.0036/M | $0.105/M | — | — | — | |
S3AI API Free | L1 98% L2 98% | qwen3.8-flash-free | free | Free | Free | — | — | — |
L1 100% | qwen3.8-flash | default | $2.40/M Cache read$0.300/M | $8.10/M | — | — | — | |
L1 100% | qwen3.8-flash | default | $38.00/M | $38.00/M | — | — | — | |
L1 99% | qwen3.8-flash | default | $4.00/M Cache read$0.400/M | $20.00/M | — | — | — | |
初叶🍂Furry API Free | L1 50% | qwen3.8-flash | free | Free | Free | — | — | — |
Alternatives & Similar Models
DeepSeek V4 Flash
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
DeepSeek V4 Pro
deepseek-v4-pro
DeepSeek V4 Pro is the professional-tier DeepSeek V4 model, targeting frontier reasoning, coding, and agent workflows with maximum capability.
GPT-5.6 Luna
gpt-5-6-luna
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
GLM 5.3 Flash
glm-5-3-flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
GPT-5.4
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
Claude Opus 4.6
claude-opus-4-6
Anthropic Claude Opus 4.6 is the most capable Claude Opus tier, optimized for complex analysis, long-horizon coding, and high-stakes enterprise reasoning workloads.
Frequently Asked Questions
- What benchmark data does Qwen3.8 Flash include?
- LMSpeed shows Qwen3.8 Flash benchmark context, API price, output speed, first-token latency, and provider data across 36 providers when those signals are available.
- What is the Qwen3.8 Flash API price?
- Qwen3.8 Flash has pricing from undefined provider} other undefined providers}}, ranging from $0.000015/M to $100.00/request. OAI2API has the lowest listed price.
- What does the Qwen3.8 Flash API pricing table include?
- The Qwen3.8 Flash API pricing table compares 36 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Qwen3.8 Flash API pricing?
- OAI2API currently has the lowest listed Qwen3.8 Flash price at $0.000015/M across undefined provider} other undefined providers}}.
- Is Qwen3.8 Flash API free?
- Yes, Qwen3.8 Flash free API options are available through 2 providerundefined other undefined} on LMSpeed, including 初叶🍂Furry API, S3AI API. These providers offer free API credits or a free tier with no per-token charges.
- Where can I get Qwen3.8 Flash free API access?
- LMSpeed currently lists 2 free API providerundefined other undefined} for Qwen3.8 Flash: 初叶🍂Furry API, S3AI API. Check each provider row before using it because free tier limits can change.
