GLM-4.1v Thinking Flash API Benchmarks, Pricing & Provider Data
Compare GLM-4.1v Thinking Flash with another model
Choose a model to open its comparison page.
GLM-4.1v Thinking Flash API pricing covers 30 API providers, from $0.0002/M to $75.00/M. The page also shows measured API speed and first-token latency.
Zhipu AI GLM-4.1v Thinking Flash is a reasoning model in the GLM series, designed for complex reasoning, problem-solving, and analytical tasks.
Specifications
Pricing Comparison
Compare GLM-4.1v Thinking Flash API pricing across 30 providers. Prices range from $0.0002/M to $75.00/M. IXIOCCAPI offers the lowest rate at $0.0002/M.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 99% | glm-4.1v-thinking-flash | default | $0.098/M | $0.098/M | — | — | — | |
L1 100% | glm-4.1v-thinking-flash | default | $75.00/M | $75.00/M | — | — | — | |
L1 100% | glm-4.1v-thinking-flash | default | $0.010/request | - | — | — | — | |
L1 99% | glm-4.1v-thinking-flash | 91vip | $0.365/M | $3.65/M | — | — | — | |
L1 99% | glm-4.1v-thinking-flash | default | $10.27/M | $10.27/M | — | — | — | |
L1 98% | zhipu/glm-4.1v-thinking-flash | 0倍倍率分组 | $0.014/M | $0.014/M | — | — | — | |
L1 100% | glm-4.1v-thinking-flash | default | $20.55/M | $20.55/M | — | — | — | |
L1 0% | glm-4.1v-thinking-flash | default | $15.41/M | $15.41/M | — | — | — | |
L1 0% | glm-4.1v-thinking-flash | default | $0.0002/M | $0.0002/M | — | — | — | |
L1 0% | glm-4.1v-thinking-flash | vip | $0.365/M | $3.65/M | — | — | — | |
L1 0% | glm-4.1v-thinking-flash | 91vip | $0.365/M | $3.65/M | — | — | — | |
L1 0% | glm-4.1v-thinking-flash | default | $10.27/M | $10.27/M | — | — | — |
Alternatives & Similar Models
GLM-4.7
glm-4-7
Zhipu GLM-4.7 is a flagship GLM release from Zhipu AI with advanced Chinese-English reasoning, coding, and agent features.
GPT-OSS
gpt-oss
GPT-OSS is an open-weight language model family designed for self-hosted inference, research, and cost-efficient alternatives to proprietary GPT-class models.
Qwen3
qwen3
Alibaba Qwen3 is the Qwen family's flagship LLM series with dense and MoE variants, seamless thinking/non-thinking modes, and leading open-source performance in math, code, and agent tasks.
GLM-4.5 Air
glm-4-5-air
Zhipu AI GLM-4.5 Air is a lightweight GLM-4.5 variant optimized for low-latency Chinese and English dialogue, retrieval-augmented apps, and edge deployments.
GLM-5
glm-5
Zhipu GLM-5 is Zhipu flagship GLM series model with enhanced reasoning, agent capabilities, and strong performance on Chinese enterprise and coding scenarios.
Gemini 2.5 Pro
gemini-2-5-pro
Google Gemini 2.5 Pro is Google advanced multimodal model with a 1M-token context window, strong STEM reasoning, and native support for images, audio, and video understanding.
Frequently Asked Questions
- What benchmark data does GLM-4.1v Thinking Flash include?
- LMSpeed shows GLM-4.1v Thinking Flash benchmark context, API price, output speed, first-token latency, and provider data across 30 providers when those signals are available.
- What is the GLM-4.1v Thinking Flash API price?
- GLM-4.1v Thinking Flash has pricing from 30 providers, ranging from $0.0002/M to $75.00/M. IXIOCCAPI has the lowest listed price.
- What does the GLM-4.1v Thinking Flash API pricing table include?
- The GLM-4.1v Thinking Flash API pricing table compares 30 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest GLM-4.1v Thinking Flash API pricing?
- IXIOCCAPI currently has the lowest listed GLM-4.1v Thinking Flash price at $0.0002/M across 30 providers.
- Can I compare GLM-4.1v Thinking Flash API price and speed together?
- Yes. LMSpeed shows GLM-4.1v Thinking Flash API price, output speed, first-token latency, and provider health on the same page so you can compare cost and performance together.
- Is GLM-4.1v Thinking Flash API free?
- GLM-4.1v Thinking Flash does not currently have a free API tier on LMSpeed. All 30 providers charge per token.
