Ling-3.0-flash API Benchmarks, Pricing & Provider Data
Compare Ling-3.0-flash with another model
Choose a model to open its comparison page.
Ling-3.0-flash benchmark, API pricing, and provider data cover undefined API provider} other undefined API providers}}, with prices starting at $0.00000001/request.
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agent...
Category Performance
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
- Coverage
- 6 / 8
- Methodology
- V3.0
#1Instruction following54.2Provisional1/4 Measured dimensions
#2Reasoning53.1Estimated2/4 Measured dimensions
#3Coding48.4Estimated3/4 Measured dimensions
#4Agents46.7Estimated2/4 Measured dimensions
#4Math46.7Estimated2/4 Measured dimensions
#6Knowledge38.9Provisional1/4 Measured dimensions
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Released
- Jul 2026
- Tokenizer
- Other
- Architecture
- text->text
- Moderated
- No
- Supported parameters
- frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstoptemperaturetool_choicetoolstop_ktop_logprobstop_p
Rankings
Excels at
It's decent at
Falls behind in
Detailed scores
Updated: Sep 11, 2026Overall
undefined metric} other undefined metrics}}
Overall score55.0#66 / 111SkillsBench44.8#2 / 3
Overall
undefined metric} other undefined metrics}}
Speed & latency
undefined metric} other undefined metrics}}
Output speed314.5 tok/s#7 / 81Time to first token1.62 s#47 / 81
Speed & latency
undefined metric} other undefined metrics}}
Pricing
undefined metric} other undefined metrics}}
Input price$0.075/M#9 / 186Output price$0.220/M#7 / 186
Pricing
undefined metric} other undefined metrics}}
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Score46.780% interval36.6–56.92/4 Measured dimensionsAgentic score51.6#51 / 76Terminal-Bench 2.157.0#10 / 10
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Coding
V3.0undefined metric} other undefined metrics}} · Estimated
Score48.480% interval39.1–57.73/4 Measured dimensionsSciCode42.0%#56 / 89Coding score44.8#66 / 87
Coding
V3.0undefined metric} other undefined metrics}} · Estimated
Reasoning
V3.0undefined metric} other undefined metrics}} · Estimated
Score53.180% interval42.3–63.92/4 Measured dimensionsGPQA85.5%#64 / 218HLE23.7%#73 / 216
Reasoning
V3.0undefined metric} other undefined metrics}} · Estimated
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Score38.980% interval24.9–52.81/4 Measured dimensionsKnowledge score34.6#81 / 83GPQA-D85.0#31 / 32
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Math
V3.0undefined metric} other undefined metrics}} · Estimated
Score46.780% interval35.1–58.22/4 Measured dimensionsMath score73.7#16 / 63AIME2693.2#13 / 14
Math
V3.0undefined metric} other undefined metrics}} · Estimated
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
Score54.280% interval38.2–70.21/4 Measured dimensionsInstruction Following score75.6#42 / 52IFBench74.5#11 / 14
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
OpenRouter endpoints
2 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
Novita novita | $0.021/M | $0.063/M | 100.0% | — | — | undefined tokens / undefined tokens |
DeepInfra deepinfra/bf16 | $0.060/M | $0.180/M | 97.2% | — | — | undefined tokens / undefined tokens |
Pricing Comparison
Compare Ling-3.0-flash API pricing across 12 providers. Prices range from $0.00000001/request to $75.00/M. Future Hub offers the lowest rate at $0.00000001/request.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% L2 100% | ling-3.0-flash | Archived | $1.00/M | $2.00/M | — | — | — | |
L1 100% L2 0% | ling-3.0-flash | default | $75.00/M | $75.00/M | — | — | — | |
L1 99% | ling-3.0-flash-free | 临时渠道2 | $0.103/M | -53%$0.103/M | — | — | — | |
L1 100% | inclusionai/ling-3.0-flash | openrouter | -92%$0.0058/M Cache read$0.0012/M | -92%$0.017/M | — | — | — | |
L1 100% | ling-3.0-flash-free | 编程 | $75.00/M | $75.00/M | — | — | — | |
L1 0% | ling-3.0-flash | openrouter | $0.00000001/request | - | — | — | — |
Alternatives & Similar Models
DeepSeek V4 Flash
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
DeepSeek V4 Pro
deepseek-v4-pro
DeepSeek V4 Pro is the professional-tier DeepSeek V4 model, targeting frontier reasoning, coding, and agent workflows with maximum capability.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
GLM-5.1
glm-5-1
Zhipu GLM-5.1 is a next-generation GLM model aimed at frontier reasoning, coding, and bilingual agent applications.
GLM-5
glm-5
Zhipu GLM-5 is Zhipu flagship GLM series model with enhanced reasoning, agent capabilities, and strong performance on Chinese enterprise and coding scenarios.
MiniMax M2.5
minimax-m2-5
MiniMax M2.5 is MiniMax's flagship text model for coding and agents, with SOTA-level programming and agentic performance, improved token efficiency, and fast high-TPS API deployment.
Frequently Asked Questions
- What benchmark data does Ling-3.0-flash include?
- LMSpeed shows Ling-3.0-flash benchmark context, API price, output speed, first-token latency, and provider data across 12 providers when those signals are available.
- What is the Ling-3.0-flash API price?
- Ling-3.0-flash has pricing from undefined provider} other undefined providers}}, ranging from $0.00000001/request to $75.00/M. Future Hub has the lowest listed price.
- What does the Ling-3.0-flash API pricing table include?
- The Ling-3.0-flash API pricing table compares 12 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Ling-3.0-flash API pricing?
- Future Hub currently has the lowest listed Ling-3.0-flash price at $0.00000001/request across undefined provider} other undefined providers}}.
- Is Ling-3.0-flash API free?
- Ling-3.0-flash does not currently have a free API tier on LMSpeed. All 12 providers charge per token.
