SpaceXAI
·Released on Mar 31, 2026

Grok 4.20 API Benchmarks, Pricing & Provider Data

Compare Grok 4.20 with another model

Choose a model to open its comparison page.

Share on X
LLM

Grok 4.20 benchmark, API pricing, and provider data cover 88 API providers, with prices starting at $0.0050/request. Grok 4.20 free API options are available from 2 providers. The page also shows measured API speed and first-token latency.

xAI Grok 4.20 is a Grok 4 series model with real-time knowledge integration, strong reasoning, and conversational capabilities for search-augmented chat.

Quality
#82of 100
45.0
LMSpeed score
Speed
75char/s
4.82 s
Cost
#120of 181
$0.0050/ 1M · 8:1 in:out
$0.0008 in · $0.0042 out

Category Performance

Observed-capability estimates with uncertainty reported separately from independent benchmark families.

Coverage
5 / 8
Methodology
V3.0
Category PerformanceObserved-capability estimates with uncertainty reported separately from independent benchmark families.Agents43.1Coding47.2Reasoning52.3Knowledge44.1Math-Multilingual-Multimodal40.2Instruction following-
#1Reasoning52.3Estimated2/4 Measured dimensions
80% interval: 41.063.5
abstract logic45.8
arc_agi · arc_agi2
scientific causal58.7
gpqa · gpqa / hle · hle / hle · hle_no_tools
multistep constraintsPrior only
evidence verificationPrior only
#2Coding47.2RatedGlobal rank #323/4 Measured dimensions
80% interval: 38.655.8
code generation54.3
livecodebench · live_code_bench_pro / scicode · scicode
repository engineering37.1
swe_pro · swe_pro / vibecode · vibe_code_bench
debugging testing50.2
swe_verified · swe_verified
tooling qualityPrior only
#3Knowledge44.1Provisional1/4 Measured dimensions
80% interval: 29.059.3
broad knowledgePrior only
professional knowledge44.1
healthbench · health_bench_hard / medxpert · med_xpert_qa_text
factualityPrior only
retrieval open bookPrior only
#4Agents43.1Estimated2/4 Measured dimensions
80% interval: 31.954.3
planning44.5
deep_planning · gert_labs
tool usePrior only
environment execution41.7
deep_search_qa · deep_search_qa / terminalbench · benchlm_agentic_terminal_bench2
recovery reliabilityPrior only
#5Multimodal40.2Estimated3/4 Measured dimensions
80% interval: 31.349.1
perception ocr44.8
simplevqa · simple_vqa
document spatial34.8
charxiv · charxiv
visual reasoning41.1
erqa · erqa / medxpert · med_xpert_qa_mm / mmmu_pro · mmmu_pro
video actionPrior only
No data:MathMultilingualInstruction following

Specifications

Input and output token limits for this model, plus how it ranks on long-context understanding.

INPUT
2Mtokens
2.4K pages of text
OUTPUT
1.8Mtokens
8K128K1M4M
2M

Features

Technical Details

Input
Output
Released
Mar 2026
Knowledge cutoff
2025-09-01
Documentation
Tokenizer
Grok
Architecture
text+image+file->text
Moderated
No
Supported parameters
include_reasoninglogprobsmax_tokensreasoningresponse_formatseedstructured_outputstemperaturetool_choicetoolstop_logprobstop_p

Rankings

Excels at

It's decent at

Falls behind in

Detailed scores

Updated: Sep 1, 2026

Overall

undefined metric} other undefined metrics}}

Overall score45.0#82 / 100
Overall score
LMSpeed rank#82
Score45.0
UpdatedAug 20, 2026
Confidence2

Pricing

undefined metric} other undefined metrics}}

Input price$1.25/M#120 / 181Output price$2.50/M#87 / 181
Input price
LMSpeed rank#120
Score$1.25/M
UpdatedSep 1, 2026
Confidence4
Output price
LMSpeed rank#87
Score$2.50/M
UpdatedSep 1, 2026
Confidence4

Agents

V3.0

undefined metric} other undefined metrics}} · Estimated

Score43.180% interval31.9–54.32/4 Measured dimensionsAgentic score30.4#60 / 69Terminal-Bench 2.047.1#47 / 53
Dimensions and evidence
Planning & decomposition44.5
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
deep_planning · gert_labs · z -0.72 · q 1.00
Tool usePrior only
Environment & long-horizon execution41.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
deep_search_qa · deep_search_qa · z -0.85 · q 1.00
terminalbench · benchlm_agentic_terminal_bench2 · z -1.00 · q 1.00
Recovery & completion reliabilityPrior only
Agentic score
LMSpeed rank#60
Score30.4
UpdatedAug 20, 2026
Confidence2
Terminal-Bench 2.0V3 evidence
LMSpeed rank#47
Score47.1
UpdatedAug 20, 2026
Confidence2
DeepSearchQAV3 evidence
LMSpeed rank#13
Score62.8
UpdatedAug 20, 2026
Confidence2
Gert LabsV3 evidence
LMSpeed rank#41
Score38.4
UpdatedAug 20, 2026
Confidence2

Coding

V3.0

undefined metric} other undefined metrics}} · Rated

Score47.2#3280% interval38.6–55.83/4 Measured dimensionsSciCode45.6%#46 / 206Coding score48.1#59 / 81
Dimensions and evidence
Code generation54.3
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
livecodebench · live_code_bench_pro · z -0.54 · q 0.50
scicode · scicode · z 0.71 · q 1.00
Repository engineering37.1
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
swe_pro · swe_pro · z -0.91 · q 1.00
vibecode · vibe_code_bench · z -1.62 · q 1.00
Debugging & testing50.2
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
swe_verified · swe_verified · z -0.02 · q 1.00
Tool-assisted development & qualityPrior only
SciCodeV3 evidence
LMSpeed rank#46
Score45.6%
UpdatedSep 1, 2026
Confidence4
Coding score
LMSpeed rank#59
Score48.1
UpdatedAug 20, 2026
Confidence2
LiveCodeBench ProV3 evidence
LMSpeed rank#3
Score74.2
UpdatedAug 20, 2026
Confidence2
SWE-bench VerifiedV3 evidence
LMSpeed rank#28
Score76.7
UpdatedAug 20, 2026
Confidence2
SWE-bench ProV3 evidence
LMSpeed rank#43
Score51.8
UpdatedAug 20, 2026
Confidence2
Vibe Code BenchV3 evidence
LMSpeed rank#30
Score4.1
UpdatedAug 20, 2026
Confidence2

Reasoning

V3.0

undefined metric} other undefined metrics}} · Estimated

Score52.380% interval41.0–63.52/4 Measured dimensionsGPQA91.1%#22 / 213HLE34.5%#40 / 210
Dimensions and evidence
Abstract logic45.8
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
arc_agi · arc_agi2 · z -0.74 · q 1.00
Scientific & causal reasoning58.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
gpqa · gpqa · z 1.07 · q 1.00
hle · hle · z 0.79 · q 1.00
hle · hle_no_tools · z -1.11 · q 1.00
Multi-step constraintsPrior only
Evidence integration & verificationPrior only
GPQAV3 evidence
LMSpeed rank#22
Score91.1%
UpdatedSep 1, 2026
Confidence4
HLEV3 evidence
LMSpeed rank#40
Score34.5%
UpdatedSep 1, 2026
Confidence4
Reasoning score
LMSpeed rank#23
Score51.2
UpdatedAug 20, 2026
Confidence2
ARC-AGI-2V3 evidence
LMSpeed rank#12
Score53.3
UpdatedAug 20, 2026
Confidence2
ARC-AGI-3
LMSpeed rank#11
Score0.1
UpdatedAug 20, 2026
Confidence2

Knowledge

V3.0

undefined metric} other undefined metrics}} · Provisional

Score44.180% interval29.0–59.31/4 Measured dimensionsGPQA-D88.5#21 / 31HLE w/o tools31.6#16 / 21
Dimensions and evidence
Broad knowledgePrior only
Professional knowledge44.1
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
healthbench · health_bench_hard · z -1.28 · q 0.88
medxpert · med_xpert_qa_text · z -0.68 · q 0.50
FactualityPrior only
Retrieval & open-book usePrior only
GPQA-D
LMSpeed rank#21
Score88.5
UpdatedAug 20, 2026
Confidence2
HLE w/o tools
LMSpeed rank#16
Score31.6
UpdatedAug 20, 2026
Confidence2
HealthBench HardV3 evidence
LMSpeed rank#6
Score20.3
UpdatedAug 20, 2026
Confidence2
MedXpertQA (Text)V3 evidence
LMSpeed rank#4
Score50.2
UpdatedAug 20, 2026
Confidence2

Math

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Foundational mathPrior only
Competition mathPrior only
Advanced proofsPrior only
Applied & tool-assisted mathPrior only

Multilingual

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Cross-language understandingPrior only
Multilingual generationPrior only
Reasoning transferPrior only
Low-resource robustnessPrior only

Multimodal

V3.0

undefined metric} other undefined metrics}} · Estimated

Score40.280% interval31.3–49.13/4 Measured dimensionsMultimodal Grounded score35.9#41 / 44MMMU-Pro75.2#25 / 31
Dimensions and evidence
Perception & OCR44.8
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
simplevqa · simple_vqa · z -0.68 · q 1.00
Document & spatial understanding34.8
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
charxiv · charxiv · z -1.90 · q 1.00
Visual reasoning41.1
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
erqa · erqa · z -1.60 · q 1.00
medxpert · med_xpert_qa_mm · z -0.83 · q 0.75
mmmu_pro · mmmu_pro · z -0.76 · q 1.00
Video & grounded actionPrior only
Multimodal Grounded score
LMSpeed rank#41
Score35.9
UpdatedAug 20, 2026
Confidence2
MMMU-ProV3 evidence
LMSpeed rank#25
Score75.2
UpdatedAug 20, 2026
Confidence2
CharXivV3 evidence
LMSpeed rank#28
Score60.9
UpdatedAug 20, 2026
Confidence2
ERQAV3 evidence
LMSpeed rank#7
Score54.1
UpdatedAug 20, 2026
Confidence2
SimpleVQAV3 evidence
LMSpeed rank#7
Score57.4
UpdatedAug 20, 2026
Confidence2
MedXpertQA (MM)V3 evidence
LMSpeed rank#5
Score65.8
UpdatedAug 20, 2026
Confidence2
Design Arena Website
LMSpeed rank#37
Score1239.0
UpdatedAug 20, 2026
Confidence2

Instruction following

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Constraint followingPrior only
Structured outputPrior only
Novel-instruction generalizationPrior only
Multi-turn & long instructionsPrior only

OpenRouter endpoints

4 endpoints

Third-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.

Provider endpointInputOutput1d uptime30m latency30m throughputContext / output
xAI
xai/priority
$2.50/M$5/M100%undefined tokens / undefined tokens
xAI
xai/zdr/priority
$2.50/M$5/M100%undefined tokens / undefined tokens
xAI
xai
$1.25/M$2.50/M99.9%undefined tokens / undefined tokens
xAI
xai/zdr
$1.25/M$2.50/M99.9%undefined tokens / undefined tokens

Pricing Comparison

Compare Grok 4.20 API pricing across 86 providers. Prices range from $0.0050/request to $75.00/M. Kunkunout API offers the lowest rate at $0.0050/request. 2 providers offer free API credits or a free tier.

ProviderHealthModel VariantGroupInput ($/M)Output ($/M)Speed (t/s)First tokenAudit
L1
98%
grok-4.20-fast
临时渠道
-98%$0.027/M
-98%$0.055/M
90.1 t/s
3.35 s
L1
98%
grok-4.20-0309
临时渠道
-98%$0.027/M
-95%$0.137/M
L1
100%
grok-4.20-fast
default
-20%$1.00/M
Cache read$0.010/M
-60%$1.00/M
L1
100%
L2
62%
grok-4.20-0309
default
$75.00/M
$75.00/M
L1
99%
grok-4.20-fast
other
-33%$0.833/M
Cache read$0.083/M
$2.50/M
L1
99%
grok-4.20-0309
other
-33%$0.833/M
Cache read$0.083/M
$2.50/M
L1
99%
grok-4.20-fast
按次计费
$0.020/request
-
L1
100%
[lq]q|YS/grok-4.20-fast
default
$6.00/request
-
L1
100%
grok-4.20-fast
default
$10.00/M
$10.00/M
L1
100%
grok-4.20-fast
default
Free
Free
L1
100%
grok-4.20-0309
default
-78%$0.274/M
$2.74/M
L1
100%
grok-4.20-fast
default
-78%$0.274/M
$2.74/M
L1
100%
grok-4.20-fast
Chat
-99%$0.014/M
-99%$0.034/M
L1
100%
grok-4.20-0309
Chat
-89%$0.137/M
-84%$0.411/M
L1
99%
grok-4.20-fast
default
-78%$0.274/M
-67%$0.822/M
L1
99%
grok-4.20-0309
default
-78%$0.274/M
-67%$0.822/M
L1
99%
grok-4.20-fast
default
-22%$0.979/M
Cache read$0.098/M
$2.94/M
L1
99%
grok-4.20-0309
default
-22%$0.979/M
Cache read$0.098/M
$2.94/M
L1
99%
grok-4.20-fast
grok
$0.0050/request
-
L1
100%
grok-4.20-beta
default
$0.010/request
-
Showing 20 model IDs of 47.

Alternatives & Similar Models

Frequently Asked Questions

What benchmark data does Grok 4.20 include?
LMSpeed shows Grok 4.20 benchmark context, API price, output speed, first-token latency, and provider data across 88 providers when those signals are available.
What is the Grok 4.20 API price?
Grok 4.20 has pricing from 88 providers, ranging from $0.0050/request to $75.00/M. Kunkunout API has the lowest listed price.
What does the Grok 4.20 API pricing table include?
The Grok 4.20 API pricing table compares 88 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
Which provider has the cheapest Grok 4.20 API pricing?
Kunkunout API currently has the lowest listed Grok 4.20 price at $0.0050/request across 88 providers.
Can I compare Grok 4.20 API price and speed together?
Yes. LMSpeed shows Grok 4.20 API price, output speed, first-token latency, and provider health on the same page so you can compare cost and performance together.
Is Grok 4.20 API free?
Yes, Grok 4.20 free API options are available through 2 providerundefined other undefined} on LMSpeed, including Moyanjdc API, Dext API. These providers offer free API credits or a free tier with no per-token charges.
Where can I get Grok 4.20 free API access?
LMSpeed currently lists 2 free API providerundefined other undefined} for Grok 4.20: Moyanjdc API, Dext API. Check each provider row before using it because free tier limits can change.

Also known as

[A-VIP]/grok-4.20-0309[lq]q|Rim/grok-4.20-beta[lq]q|XJ/grok-4.20-fast[lq]q|YS/grok-4.20-0309[lq]q|YS/grok-4.20-fast

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.Standard benchmark data may include BenchLM and other public sources.

Build with LMSpeed Data

Free

Free public API for LLM pricing, benchmarks & provider data

  • Real-time pricing & availability
  • Speed & latency benchmarks
  • 300+ models, 600+ providers
  • 1,000 requests/day
View API Documentation