SpaceXAI
·Released on Mar 31, 2026

Grok 4.20 API Benchmarks, Pricing & Provider Data

Compare Grok 4.20 with another model

Choose a model to open its comparison page.

Share on X
LLM

Grok 4.20 benchmark, API pricing, and provider data cover undefined API provider} other undefined API providers}}, with prices starting at $0.0068/request. The page also shows measured API speed and first-token latency.

xAI Grok 4.20 is a Grok 4 series model with real-time knowledge integration, strong reasoning, and conversational capabilities for search-augmented chat.

Quality
#95of 112
45.0
LMSpeed score
Speed
75char/s
4.82 s
Cost
#123of 186
$0.0068/ 1M · 8:1 in:out
$0.0011 in · $0.0058 out

Category Performance

Observed-capability estimates with uncertainty reported separately from independent benchmark families.

Coverage
5 / 8
Methodology
V3.0
Category PerformanceObserved-capability estimates with uncertainty reported separately from independent benchmark families.Agents42.6Coding45Reasoning51.4Knowledge43.4Math-Multilingual-Multimodal40.1Instruction following-
#1Reasoning51.4Estimated2/4 Measured dimensions
80% interval: 40.262.6
abstract logic44.8
arc_agi · arc_agi2
scientific causal58.1
gpqa · gpqa / hle · hle / hle · hle_no_tools
multistep constraintsPrior only
evidence verificationPrior only
#2Coding45RatedGlobal rank #363/4 Measured dimensions
80% interval: 35.954.1
code generation48.4
livecodebench · live_code_bench_pro
repository engineering36.7
swe_pro · swe_pro / vibecode · vibe_code_bench
debugging testing49.9
swe_verified · swe_verified
tooling qualityPrior only
#3Knowledge43.4Provisional1/4 Measured dimensions
80% interval: 28.358.4
broad knowledgePrior only
professional knowledge43.4
healthbench · health_bench_hard / medxpert · med_xpert_qa_text
factualityPrior only
retrieval open bookPrior only
#4Agents42.6Estimated2/4 Measured dimensions
80% interval: 31.453.8
planning44.5
deep_planning · gert_labs
tool usePrior only
environment execution40.8
deep_search_qa · deep_search_qa / terminalbench · benchlm_agentic_terminal_bench2
recovery reliabilityPrior only
#5Multimodal40.1Estimated3/4 Measured dimensions
80% interval: 31.249.0
perception ocr44.8
simplevqa · simple_vqa
document spatial34.6
charxiv · charxiv
visual reasoning41.1
erqa · erqa / medxpert · med_xpert_qa_mm / mmmu_pro · mmmu_pro
video actionPrior only
No data:MathMultilingualInstruction following

Specifications

Input and output token limits for this model, plus how it ranks on long-context understanding.

INPUT
2Mtokens
2.4K pages of text
OUTPUT
1.8Mtokens
8K128K1M4M
2M

Features

Technical Details

Input
Output
Released
Mar 2026
Knowledge cutoff
2025-09-01
Documentation
Tokenizer
Grok
Architecture
text+image+file->text
Moderated
No
Supported parameters
include_reasoninglogprobsmax_tokensreasoningresponse_formatseedstructured_outputstemperaturetool_choicetoolstop_logprobstop_p

Rankings

Excels at

It's decent at

Falls behind in

Detailed scores

Updated: Sep 13, 2026

Overall

undefined metric} other undefined metrics}}

Overall score45.0#95 / 112
Overall score
LMSpeed rank#95
Score45.0
UpdatedAug 20, 2026
Confidence2

Pricing

undefined metric} other undefined metrics}}

Input price$1.25/M#123 / 186Output price$2.50/M#89 / 186
Input price
LMSpeed rank#123
Score$1.25/M
UpdatedSep 13, 2026
Confidence4
Output price
LMSpeed rank#89
Score$2.50/M
UpdatedSep 13, 2026
Confidence4

Agents

V3.0

undefined metric} other undefined metrics}} · Estimated

Score42.680% interval31.4–53.82/4 Measured dimensionsAgentic score30.4#68 / 77Terminal-Bench 2.047.1#47 / 53
Dimensions and evidence
Planning & decomposition44.5
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
deep_planning · gert_labs · z -0.72 · q 1.00
Tool usePrior only
Environment & long-horizon execution40.8
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
deep_search_qa · deep_search_qa · z -0.97 · q 1.00
terminalbench · benchlm_agentic_terminal_bench2 · z -1.00 · q 1.00
Recovery & completion reliabilityPrior only
Agentic score
LMSpeed rank#68
Score30.4
UpdatedAug 20, 2026
Confidence2
Terminal-Bench 2.0V3 evidence
LMSpeed rank#47
Score47.1
UpdatedAug 20, 2026
Confidence2
DeepSearchQAV3 evidence
LMSpeed rank#14
Score62.8
UpdatedAug 20, 2026
Confidence2
Gert LabsV3 evidence
LMSpeed rank#41
Score38.4
UpdatedAug 20, 2026
Confidence2

Coding

V3.0

undefined metric} other undefined metrics}} · Rated

Score45#3680% interval35.9–54.13/4 Measured dimensionsCoding score48.1#61 / 87LiveCodeBench Pro74.2#3 / 4
Dimensions and evidence
Code generation48.4
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
livecodebench · live_code_bench_pro · z -0.54 · q 0.50
Repository engineering36.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
swe_pro · swe_pro · z -0.92 · q 1.00
vibecode · vibe_code_bench · z -1.62 · q 1.00
Debugging & testing49.9
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
swe_verified · swe_verified · z -0.02 · q 1.00
Tool-assisted development & qualityPrior only
Coding score
LMSpeed rank#61
Score48.1
UpdatedAug 20, 2026
Confidence2
LiveCodeBench ProV3 evidence
LMSpeed rank#3
Score74.2
UpdatedAug 20, 2026
Confidence2
SWE-bench VerifiedV3 evidence
LMSpeed rank#28
Score76.7
UpdatedAug 20, 2026
Confidence2
SWE-bench ProV3 evidence
LMSpeed rank#44
Score51.8
UpdatedAug 20, 2026
Confidence2
Vibe Code BenchV3 evidence
LMSpeed rank#30
Score4.1
UpdatedAug 20, 2026
Confidence2

Reasoning

V3.0

undefined metric} other undefined metrics}} · Estimated

Score51.480% interval40.2–62.62/4 Measured dimensionsGPQA91.1%#25 / 218HLE34.5%#45 / 216
Dimensions and evidence
Abstract logic44.8
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
arc_agi · arc_agi2 · z -0.78 · q 1.00
Scientific & causal reasoning58.1
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
gpqa · gpqa · z 1.00 · q 1.00
hle · hle · z 0.71 · q 1.00
hle · hle_no_tools · z -1.30 · q 1.00
Multi-step constraintsPrior only
Evidence integration & verificationPrior only
GPQAV3 evidence
LMSpeed rank#25
Score91.1%
UpdatedSep 13, 2026
Confidence4
HLEV3 evidence
LMSpeed rank#45
Score34.5%
UpdatedSep 13, 2026
Confidence4
Reasoning score
LMSpeed rank#58
Score51.2
UpdatedAug 20, 2026
Confidence2
ARC-AGI-2V3 evidence
LMSpeed rank#14
Score53.3
UpdatedAug 20, 2026
Confidence2
ARC-AGI-3
LMSpeed rank#12
Score0.1
UpdatedAug 20, 2026
Confidence2

Knowledge

V3.0

undefined metric} other undefined metrics}} · Provisional

Score43.480% interval28.3–58.41/4 Measured dimensionsGPQA-D88.5#22 / 32HLE w/o tools31.6#17 / 22
Dimensions and evidence
Broad knowledgePrior only
Professional knowledge43.4
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
healthbench · health_bench_hard · z -1.24 · q 1.00
medxpert · med_xpert_qa_text · z -0.68 · q 0.50
FactualityPrior only
Retrieval & open-book usePrior only
GPQA-D
LMSpeed rank#22
Score88.5
UpdatedAug 20, 2026
Confidence2
HLE w/o tools
LMSpeed rank#17
Score31.6
UpdatedAug 20, 2026
Confidence2
HealthBench HardV3 evidence
LMSpeed rank#7
Score20.3
UpdatedAug 20, 2026
Confidence2
MedXpertQA (Text)V3 evidence
LMSpeed rank#4
Score50.2
UpdatedAug 20, 2026
Confidence2

Math

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Foundational mathPrior only
Competition mathPrior only
Advanced proofsPrior only
Applied & tool-assisted mathPrior only

Multilingual

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Cross-language understandingPrior only
Multilingual generationPrior only
Reasoning transferPrior only
Low-resource robustnessPrior only

Multimodal

V3.0

undefined metric} other undefined metrics}} · Estimated

Score40.180% interval31.2–49.03/4 Measured dimensionsMultimodal Grounded score35.9#53 / 56MMMU-Pro75.2#25 / 31
Dimensions and evidence
Perception & OCR44.8
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
simplevqa · simple_vqa · z -0.68 · q 1.00
Document & spatial understanding34.6
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
charxiv · charxiv · z -1.91 · q 1.00
Visual reasoning41.1
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
erqa · erqa · z -1.60 · q 1.00
medxpert · med_xpert_qa_mm · z -0.83 · q 0.75
mmmu_pro · mmmu_pro · z -0.76 · q 1.00
Video & grounded actionPrior only
Multimodal Grounded score
LMSpeed rank#53
Score35.9
UpdatedAug 20, 2026
Confidence2
MMMU-ProV3 evidence
LMSpeed rank#25
Score75.2
UpdatedAug 20, 2026
Confidence2
CharXivV3 evidence
LMSpeed rank#29
Score60.9
UpdatedAug 20, 2026
Confidence2
ERQAV3 evidence
LMSpeed rank#7
Score54.1
UpdatedAug 20, 2026
Confidence2
SimpleVQAV3 evidence
LMSpeed rank#7
Score57.4
UpdatedAug 20, 2026
Confidence2
MedXpertQA (MM)V3 evidence
LMSpeed rank#5
Score65.8
UpdatedAug 20, 2026
Confidence2
Design Arena Website
LMSpeed rank#40
Score1239.0
UpdatedAug 20, 2026
Confidence2

Instruction following

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Constraint followingPrior only
Structured outputPrior only
Novel-instruction generalizationPrior only
Multi-turn & long instructionsPrior only

OpenRouter endpoints

4 endpoints

Third-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.

Provider endpointInputOutput1d uptime30m latency30m throughputContext / output
xAI
xai/zdr/priority
$2.50/M$5/M100%undefined tokens / undefined tokens
xAI
xai
$1.25/M$2.50/M99.9%undefined tokens / undefined tokens
xAI
xai/zdr
$1.25/M$2.50/M99.8%undefined tokens / undefined tokens
xAI
xai/priority
$2.50/M$5/M99.7%undefined tokens / undefined tokens

Pricing Comparison

Compare Grok 4.20 API pricing across 51 providers. Prices range from $0.0068/request to $1071.43/M. YX 公益站 offers the lowest rate at $0.0068/request.

ProviderHealthModel VariantGroupInput ($/M)Output ($/M)Speed (t/s)First tokenAudit
L1
99%
grok-4.20-fast
临时渠道
-98%$0.027/M
-98%$0.055/M
90.1 t/s
3.35 s
L1
99%
grok-4.20-0309
临时渠道
-98%$0.027/M
-95%$0.137/M
L1
100%
grok-4.20-fast
Chat
-99%$0.014/M
-99%$0.034/M
L1
100%
grok-4.20-0309
grok
-89%$0.137/M
-84%$0.411/M
L1
100%
grok-4.20-fast
default
-78%$0.274/M
$2.74/M
L1
100%
grok-4.20-fast
default
-20%$1.00/M
Cache read$0.010/M
-60%$1.00/M
L1
100%
[lq]q|YS/grok-4.20-fast
default
$4.00/request
-
L1
100%
L2
0%
grok-4.20-0309
default
$75.00/M
$75.00/M
L1
100%
grok-4.20-fast
default
-78%$0.274/M
Cache read$0.027/M
-67%$0.822/M
L1
99%
grok-4.20-fast
default
$1071.43/M
$1071.43/M
L1
0%
grok-4.20-fast
2api
-96%$0.050/M
-98%$0.050/M
L1
0%
grok-4.20
懒人
$1.50/M
Cache read$0.100/MCache write$3.00/MCache write 1h$4.80/M
-80%$0.500/M
L1
61%
grok-4.20-beta
default
$0.010/request
-
L1
100%
grok-4.20-fast
default
$0.021/request
-
L1
100%
grok-4.20-0309
default
-70%$0.371/M
Cache read$0.059/M
-70%$0.743/M
L1
100%
grok-4.20-fast
default
$11.14/M
$11.14/M
L1
100%
grok-4.20-fast
公司
$75.00/M
$75.00/M
L1
100%
grok-4.20-0309
公司
$75.00/M
$75.00/M
L1
99%
grok-4.20
default
$37.50/M
$37.50/M
L1
0%
grok-4.20-0309
default
$0.014/request
-
Showing 20 model IDs of 30.

Alternatives & Similar Models

Frequently Asked Questions

What benchmark data does Grok 4.20 include?
LMSpeed shows Grok 4.20 benchmark context, API price, output speed, first-token latency, and provider data across 51 providers when those signals are available.
What is the Grok 4.20 API price?
Grok 4.20 has pricing from undefined provider} other undefined providers}}, ranging from $0.0068/request to $1071.43/M. YX 公益站 has the lowest listed price.
What does the Grok 4.20 API pricing table include?
The Grok 4.20 API pricing table compares 51 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
Which provider has the cheapest Grok 4.20 API pricing?
YX 公益站 currently has the lowest listed Grok 4.20 price at $0.0068/request across undefined provider} other undefined providers}}.
Can I compare Grok 4.20 API price and speed together?
Yes. LMSpeed shows Grok 4.20 API price, output speed, first-token latency, and provider health on the same page so you can compare cost and performance together.
Is Grok 4.20 API free?
Grok 4.20 does not currently have a free API tier on LMSpeed. All 51 providers charge per token.

Also known as

[A-VIP]/grok-4.20-0309[lq]q|Rim/grok-4.20-beta[lq]q|XJ/grok-4.20-fast[lq]q|YS/grok-4.20-0309[lq]q|YS/grok-4.20-fast

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.Standard benchmark data may include BenchLM and other public sources.

Build with LMSpeed Data

Free

Free public API for LLM pricing, benchmarks & provider data

  • Real-time pricing & availability
  • Speed & latency benchmarks
  • 300+ models, 600+ providers
  • 1,000 requests/day
View API Documentation