GLM 5.3 Flash API Benchmarks, Pricing & Provider Data
Compare GLM 5.3 Flash with another model
Choose a model to open its comparison page.
GLM 5.3 Flash benchmark, API pricing, and provider data cover undefined API provider} other undefined API providers}}, with prices starting at $0.000015/M. GLM 5.3 Flash free API options are available from undefined provider} other undefined providers}}. The page also shows measured API speed and first-token latency.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-contex...
Category Performance
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
- Coverage
- 5 / 8
- Methodology
- V3.0
#1Agents60.6Estimated2/4 Measured dimensions
#2Multimodal57.9Provisional1/4 Measured dimensions
#3Reasoning57.6Estimated2/4 Measured dimensions
#4Coding56.6Estimated3/4 Measured dimensions
#5Knowledge54.2Provisional1/4 Measured dimensions
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Released
- Aug 2026
- Tokenizer
- Other
- Architecture
- text+image+video->text
- Moderated
- No
- Expiration date
- 2098-12-31
- Supported parameters
- frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Rankings
Excels at
It's decent at
Falls behind in
Detailed scores
Updated: Sep 14, 2026Overall
undefined metric} other undefined metrics}}
Overall score67.0#27 / 112DeepSWE63.4#12 / 21
Overall
undefined metric} other undefined metrics}}
Speed & latency
undefined metric} other undefined metrics}}
Output speed114.1 tok/s#38 / 80Time to first token2.11 s#54 / 80
Speed & latency
undefined metric} other undefined metrics}}
Pricing
undefined metric} other undefined metrics}}
Input price$0.150/M#22 / 186Output price$0.500/M#22 / 186
Pricing
undefined metric} other undefined metrics}}
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Score60.680% interval48.7–72.52/4 Measured dimensionsAgentic score71.1#30 / 77Terminal-Bench 2.184.3#6 / 10
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Coding
V3.0undefined metric} other undefined metrics}} · Estimated
Score56.680% interval47.3–65.93/4 Measured dimensionsSciCode51.6%#27 / 89Coding score55.4#38 / 87
Coding
V3.0undefined metric} other undefined metrics}} · Estimated
Reasoning
V3.0undefined metric} other undefined metrics}} · Estimated
Score57.680% interval46.4–68.72/4 Measured dimensionsGPQA91.2%#24 / 218HLE39.9%#31 / 216
Reasoning
V3.0undefined metric} other undefined metrics}} · Estimated
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Score54.280% interval40.2–68.21/4 Measured dimensionsKnowledge score71.8#35 / 83Artificial Analysis Intelligence Index57.5#3 / 117
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Math
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Math
V3.0undefined metric} other undefined metrics}} · No data
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
Multimodal
V3.0undefined metric} other undefined metrics}} · Provisional
Score57.980% interval41.7–74.01/4 Measured dimensionsMultimodal Grounded score80.5#12 / 56OfficeQA Pro62.4#5 / 10
Multimodal
V3.0undefined metric} other undefined metrics}} · Provisional
Instruction following
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Instruction following
V3.0undefined metric} other undefined metrics}} · No data
OpenRouter endpoints
26 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
BaseTen baseten/fp8 | $0.150/M | $0.500/M | 99.9% | — | — | undefined tokens / undefined tokens |
CoreWeave coreweave/fp8 | $0.150/M | $0.500/M | 99.8% | — | — | undefined tokens / undefined tokens |
Relace relace | $0.090/M | $0.300/M | 99.8% | — | — | undefined tokens / undefined tokens |
NextBit nextbit/fp8 | $0.177/M | $0.590/M | 99.6% | — | — | undefined tokens / undefined tokens |
Sail Research sail-research/fp8 | $0.150/M | $0.500/M | 99.6% | — | — | undefined tokens / undefined tokens |
Cloudflare cloudflare | $0.150/M | $0.500/M | 99.4% | — | — | undefined tokens / undefined tokens |
Reka reka/fp8 | $0.132/M | $0.440/M | 99.2% | — | — | undefined tokens / undefined tokens |
Modal modal/fp8 | $0.450/M | $1.50/M | 99.1% | — | — | undefined tokens / undefined tokens |
DigitalOcean digitalocean | $0.150/M | $0.500/M | 99.1% | — | — | undefined tokens / undefined tokens |
Together together | $0.150/M | $0.500/M | 99.1% | — | — | undefined tokens / undefined tokens |
StreamLake streamlake/fp8 | $0.112/M | $0.374/M | 99.0% | — | — | undefined tokens / undefined tokens |
Phala phala/fp8 | $0.150/M | $0.500/M | 99.0% | — | — | undefined tokens / undefined tokens |
Venice venice | $0.150/M | $0.500/M | 98.9% | — | — | undefined tokens / undefined tokens |
Wafer wafer | $0.100/M | $0.350/M | 98.9% | — | — | undefined tokens / undefined tokens |
DeepInfra deepinfra/fp4 | $0.075/M | $0.250/M | 98.9% | — | — | undefined tokens / undefined tokens |
Novita novita/fp8 | $0.132/M | $0.440/M | 98.9% | — | — | undefined tokens / undefined tokens |
GMICloud gmicloud/fp8 | $0.112/M | $0.375/M | 98.8% | — | — | undefined tokens / undefined tokens |
Parasail parasail/fp8 | $0.150/M | $0.500/M | 98.6% | — | — | undefined tokens / undefined tokens |
Io Net io-net/fp8 | $0.150/M | $0.500/M | 98.2% | — | — | undefined tokens / undefined tokens |
SiliconFlow siliconflow/fp8 | $0.150/M | $0.500/M | 97.6% | — | — | undefined tokens / undefined tokens |
Z.AI z-ai/fp8 | $0.150/M | $0.500/M | 97.6% | — | — | undefined tokens / undefined tokens |
Friendli friendli | $0.150/M | $0.500/M | 94.3% | — | — | undefined tokens / undefined tokens |
Crusoe crusoe/fp4 | $0.150/M | $0.500/M | 92.9% | — | — | undefined tokens / undefined tokens |
Fireworks fireworks | $0.150/M | $0.500/M | 91.9% | — | — | undefined tokens / undefined tokens |
Makora makora | $0.140/M | $0.470/M | 87.8% | — | — | undefined tokens / undefined tokens |
Morph morph/fp8 | $0.100/M | $0.350/M | 83.1% | — | — | undefined tokens / undefined tokens |
Pricing Comparison
Compare GLM 5.3 Flash API pricing across 131 providers. Prices range from $0.000015/M to $547.50/M. OAI2API offers the lowest rate at $0.000015/M. 9 providers offer free API credits or a free tier.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% | glm-5.3-flash | sale | $30.00/M | $30.00/M | — | — | — | |
L1 100% | glm-5.3-flash | zai-officially | -24%$0.114/M Cache read$0.033/M | -20%$0.400/M | — | — | — | |
L1 99% L2 0% | z-ai/glm-5.3-flash | default | $0.150/M Cache read$0.030/M | $0.500/M | — | — | — | |
L1 100% | glm-5.3-flash | default | -63%$0.055/M | -62%$0.192/M | — | — | — | |
L1 100% | glm-5.3-flash | default | -63%$0.055/M Cache read$0.016/M | -62%$0.192/M | — | — | — | |
L1 100% | glm-5.3-flash | Self-Deployed-2 | -85%$0.022/M Cache read$0.0044/M | -85%$0.074/M | — | — | — | |
L1 100% | glm-5.3-flash | deepseek | -71%$0.044/M | -69%$0.153/M | — | — | — | |
L1 100% | glm-5.3-flash | default | -57%$0.064/M Cache read$0.085/M | -57%$0.213/M | — | — | — | |
L1 100% L2 100% | glm-5.3-flash | OpenModels | -47%$0.080/M Cache read$0.023/MCache write$0.080/MCache write 1h$0.128/M | -44%$0.280/M | — | — | — | |
L1 100% | glm-5.3-flash | default | $0.400/M | $1.40/M | — | — | — | |
L1 100% | glm-5.3-flash | glm | $0.280/M Cache read$0.081/M | $0.980/M | — | — | — | |
L1 100% | glm-5.3-flash | glm | $0.480/M Cache read$0.138/M | $1.68/M | — | — | — | |
L1 100% | [hm]q|满血/glm-5.3-flash | default | $10.00/request | - | — | — | — | |
DeadlySignal API Free | L1 99% L2 100% | z-ai/glm-5.3-flash | nvidia-free | Free | Free | — | — | — |
L1 99% L2 100% | glm-5.3-flash | diamond-glm | $54.75/M | $54.75/M | — | — | — | |
L1 100% | glm-5.3-flash | GLM-免费科研🔥-x0.01-福利 | -99%$0.0015/M Cache read$0.0003/M | -99%$0.0050/M | — | — | — | |
L1 100% | glm-5.3-flash | default | $0.036/request | - | — | — | — | |
L1 100% L2 0% | glm-5.3-flash | 智谱 国模折扣 | -62%$0.058/M Cache read$0.014/M | -60%$0.200/M | — | — | — | |
L1 100% | z-ai/glm-5.3-flash | default | -50%$0.075/M Cache read$0.015/M | -50%$0.250/M | — | — | — | |
L1 100% | glm-5.3-flash | default | -40%$0.090/M Cache read$0.018/M | -40%$0.300/M | — | — | — |
Alternatives & Similar Models
DeepSeek V4 Flash
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
DeepSeek V4 Pro
deepseek-v4-pro
DeepSeek V4 Pro is the professional-tier DeepSeek V4 model, targeting frontier reasoning, coding, and agent workflows with maximum capability.
GLM-5.2
glm-5-2
Zhipu GLM-5.2 is Zhipu latest flagship coding and agentic model with a 1M-token context window, enhanced reasoning modes, and long-horizon software engineering capabilities.
GPT-5.6 Luna
gpt-5-6-luna
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
Kimi K3
kimi-k3
Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating...
GPT-5.4
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
Frequently Asked Questions
- What benchmark data does GLM 5.3 Flash include?
- LMSpeed shows GLM 5.3 Flash benchmark context, API price, output speed, first-token latency, and provider data across 140 providers when those signals are available.
- What is the GLM 5.3 Flash API price?
- GLM 5.3 Flash has pricing from undefined provider} other undefined providers}}, ranging from $0.000015/M to $547.50/M. OAI2API has the lowest listed price.
- What does the GLM 5.3 Flash API pricing table include?
- The GLM 5.3 Flash API pricing table compares 140 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest GLM 5.3 Flash API pricing?
- OAI2API currently has the lowest listed GLM 5.3 Flash price at $0.000015/M across undefined provider} other undefined providers}}.
- Can I compare GLM 5.3 Flash API price and speed together?
- Yes. LMSpeed shows GLM 5.3 Flash API price, output speed, first-token latency, and provider health on the same page so you can compare cost and performance together.
- Is GLM 5.3 Flash API free?
- Yes, GLM 5.3 Flash free API options are available through 9 providerundefined other undefined} on LMSpeed, including 兔子API, 兔子API, DeadlySignal API, DeadlySignal API, S3AI API. These providers offer free API credits or a free tier with no per-token charges.
- Where can I get GLM 5.3 Flash free API access?
- LMSpeed currently lists 9 free API providerundefined other undefined} for GLM 5.3 Flash: 兔子API, 兔子API, DeadlySignal API, DeadlySignal API, S3AI API. Check each provider row before using it because free tier limits can change.
