GLM 5.3 API Benchmarks, Pricing & Provider Data
Compare GLM 5.3 with another model
Choose a model to open its comparison page.
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves....
Category Performance
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
- Coverage
- 2 / 8
- Methodology
- V3.0
#1Coding63.3Provisional1/4 Measured dimensions
#2Reasoning61.2Provisional1/4 Measured dimensions
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Rankings
Excels at
Detailed scores
Updated: Aug 20, 2026Speed & latency
undefined metric} other undefined metrics}}
Output speed79.6 tok/s#42 / 73Time to first token1.51 s#45 / 73
Speed & latency
undefined metric} other undefined metrics}}
Pricing
undefined metric} other undefined metrics}}
Input price$1.40/M#132 / 180Output price$4.40/M#115 / 180
Pricing
undefined metric} other undefined metrics}}
Agents
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Agents
V3.0undefined metric} other undefined metrics}} · No data
Coding
V3.0undefined metric} other undefined metrics}} · Provisional
Score63.380% interval47.3–79.21/4 Measured dimensionsSciCode56.5%#7 / 204
Coding
V3.0undefined metric} other undefined metrics}} · Provisional
Reasoning
V3.0undefined metric} other undefined metrics}} · Provisional
Score61.280% interval47.3–75.21/4 Measured dimensionsGPQA91.7%#18 / 211HLE42.3%#21 / 208
Reasoning
V3.0undefined metric} other undefined metrics}} · Provisional
Knowledge
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Knowledge
V3.0undefined metric} other undefined metrics}} · No data
Math
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Math
V3.0undefined metric} other undefined metrics}} · No data
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
Instruction following
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Instruction following
V3.0undefined metric} other undefined metrics}} · No data
OpenRouter endpoints
1 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
Z.AI z-ai/fp8 | $1.40/M | $4.40/M | 100.0% | — | — | undefined tokens / undefined tokens |
Alternatives & Similar Models
DeepSeek V4 Flash
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
DeepSeek V4 Pro
deepseek-v4-pro
DeepSeek V4 Pro is the professional-tier DeepSeek V4 model, targeting frontier reasoning, coding, and agent workflows with maximum capability.
GLM-5.1
glm-5-1
Zhipu GLM-5.1 is a next-generation GLM model aimed at frontier reasoning, coding, and bilingual agent applications.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
GLM-5
glm-5
Zhipu GLM-5 is Zhipu flagship GLM series model with enhanced reasoning, agent capabilities, and strong performance on Chinese enterprise and coding scenarios.
MiniMax M2.5
minimax-m2-5
MiniMax M2.5 is MiniMax's flagship text model for coding and agents, with SOTA-level programming and agentic performance, improved token efficiency, and fast high-TPS API deployment.
