OpenAI
·Released on Mar 5, 2026

GPT-5.4 API Benchmarks, Pricing & Provider Data

Compare GPT-5.4 with another model

Choose a model to open its comparison page.

Share on X
LLM

GPT-5.4 benchmark, API pricing, and provider data cover 1010 API providers, with prices starting at $0.0027/request. GPT-5.4 free API options are available from 10 providers. The page also shows measured API speed and first-token latency.

OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.

Quality
#29of 101
64.0
LMSpeed score
Speed
49char/s
4.45 s
Cost
#151of 182
$0.0027/ 1M · 8:1 in:out
$0.0004 in · $0.0023 out

Category Performance

Observed-capability estimates with uncertainty reported separately from independent benchmark families.

Coverage
7 / 8
Methodology
V3.0
Category PerformanceObserved-capability estimates with uncertainty reported separately from independent benchmark families.Agents57Coding62.2Reasoning55.9Knowledge59.9Math57.6Multilingual-Multimodal52.9Instruction following53.9
#1Coding62.2Estimated2/4 Measured dimensions
80% interval: 51.972.5
code generation65.2
livecodebench · live_code_bench_pro / scicode · scicode
repository engineering59.2
react_native_bench · react_native_evals / swe_pro · swe_pro / vibecode · vibe_code_bench
debugging testingPrior only
tooling qualityPrior only
#2Knowledge59.9Provisional1/4 Measured dimensions
80% interval: 44.775.0
broad knowledgePrior only
professional knowledge59.9
healthbench · health_bench_hard / medxpert · med_xpert_qa_text
factualityPrior only
retrieval open bookPrior only
#3Math57.6Provisional1/4 Measured dimensions
80% interval: 41.573.6
foundational mathPrior only
competition mathPrior only
proof frontier57.6
frontiermath · frontier_math_v2_tier4 / frontiermath · frontier_math_v2_tiers13
applied tool mathPrior only
#4Agents57RatedGlobal rank #174/4 Measured dimensions
80% interval: 51.962.1
planning58.3
deep_planning · gert_labs
tool use60.2
mcp_atlas · mcp_atlas / tau · tau2_bench / toolathlon · toolathlon
environment execution56
browsecomp · browse_comp / deep_search_qa · deep_search_qa / gdpval_aa · benchlm_agentic_gdpval_aa / osworld · os_world_verified / terminalbench · benchlm_agentic_terminal_bench2
recovery reliability53.6
apex_agents · apex_agents_aa / claw_eval · claw_eval / cyber_gym · cyber_gym / exploit_gym · exploit_gym / jobbench · job_bench / researchclaw · research_claw_bench
#5Reasoning55.9RatedGlobal rank #153/4 Measured dimensions
80% interval: 47.264.6
abstract logic51.4
arc_agi · arc_agi2
scientific causal59.6
critpt · critpt / gpqa · gpqa / hle · hle / hle · hle_no_tools
multistep constraintsPrior only
evidence verification56.7
lcr · lcr
#6Instruction following53.9Provisional1/4 Measured dimensions
80% interval: 37.969.9
constraint followingPrior only
structured outputPrior only
novel instruction generalization53.9
ifbench · aa_if_bench
long multiturn instructionPrior only
#7Multimodal52.9RatedGlobal rank #44/4 Measured dimensions
80% interval: 46.259.5
perception ocr46.7
simplevqa · simple_vqa
document spatial50.4
charxiv · charxiv
visual reasoning56.9
erqa · erqa / medxpert · med_xpert_qa_mm / mmmu_pro · mmmu_pro
video action57.6
screenspot · screen_spot_pro
No data:Multilingual

Specifications

Input and output token limits for this model, plus how it ranks on long-context understanding.

INPUT
1.1Mtokens
1.3K pages of text
OUTPUT
128Ktokens
8K128K1M4M
1.1M

Features

Technical Details

Input
Output
Released
Mar 2026
Documentation
Tokenizer
GPT
Architecture
text+image+file->text
Moderated
Yes
Supported parameters
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Rankings

Excels at

It's decent at

Falls behind in

Detailed scores

Updated: Sep 2, 2026

Overall

undefined metric} other undefined metrics}}

Overall score64.0#29 / 101
Overall score
LMSpeed rank#29
Score64.0
UpdatedSep 2, 2026
Confidence3

Pricing

undefined metric} other undefined metrics}}

Input price$2.50/M#151 / 182Output price$15.00/M#156 / 182
Input price
LMSpeed rank#151
Score$2.50/M
UpdatedSep 2, 2026
Confidence4
Output price
LMSpeed rank#156
Score$15.00/M
UpdatedSep 2, 2026
Confidence4

Agents

V3.0

undefined metric} other undefined metrics}} · Rated

Score57#1780% interval51.9–62.14/4 Measured dimensionsAgentic score74.8#23 / 70Terminal-Bench 2.075.1#13 / 53
Dimensions and evidence
Planning & decomposition58.3
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
deep_planning · gert_labs · z 1.12 · q 1.00
Tool use60.2
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
mcp_atlas · mcp_atlas · z -0.23 · q 1.00
tau · tau2_bench · z 1.46 · q 1.00
toolathlon · toolathlon · z 0.67 · q 1.00
Environment & long-horizon execution56
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
browsecomp · browse_comp · z 0.21 · q 1.00
deep_search_qa · deep_search_qa · z -0.41 · q 1.00
gdpval_aa · benchlm_agentic_gdpval_aa · z 0.60 · q 1.00
osworld · os_world_verified · z 0.18 · q 1.00
terminalbench · benchlm_agentic_terminal_bench2 · z 0.79 · q 1.00
Recovery & completion reliability53.6
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
apex_agents · apex_agents_aa · z 0.51 · q 1.00
claw_eval · claw_eval · z 0.11 · q 1.00
cyber_gym · cyber_gym · z 0.23 · q 1.00
exploit_gym · exploit_gym · z -0.95 · q 0.50
jobbench · job_bench · z 0.39 · q 1.00
researchclaw · research_claw_bench · z -0.89 · q 1.00
Agentic score
LMSpeed rank#23
Score74.8
UpdatedSep 2, 2026
Confidence3
Terminal-Bench 2.0V3 evidence
LMSpeed rank#13
Score75.1
UpdatedSep 2, 2026
Confidence3
CyberGymV3 evidence
LMSpeed rank#6
Score79.0
UpdatedSep 2, 2026
Confidence3
BrowseCompV3 evidence
LMSpeed rank#15
Score82.7
UpdatedSep 2, 2026
Confidence3
OSWorld-VerifiedV3 evidence
LMSpeed rank#11
Score75.0
UpdatedSep 2, 2026
Confidence3
MCP AtlasV3 evidence
LMSpeed rank#19
Score70.6
UpdatedSep 2, 2026
Confidence3
ToolathlonV3 evidence
LMSpeed rank#6
Score54.6
UpdatedSep 2, 2026
Confidence3
Τ²-bench resultsV3 evidence
LMSpeed rank#2
Score98.9
UpdatedSep 2, 2026
Confidence3
Claw-EvalV3 evidence
LMSpeed rank#13
Score60.3
UpdatedSep 2, 2026
Confidence3
DeepSearchQAV3 evidence
LMSpeed rank#11
Score73.6
UpdatedSep 2, 2026
Confidence3
AA Agentic Index
LMSpeed rank#20
Score44.2
UpdatedSep 2, 2026
Confidence3
APEX-Agents-AAV3 evidence
LMSpeed rank#7
Score33.3
UpdatedSep 2, 2026
Confidence3
GDPval-AA
LMSpeed rank#21
Score44.2
UpdatedSep 2, 2026
Confidence3
GDPval-AAV3 evidence
LMSpeed rank#19
Score1385.0
UpdatedSep 2, 2026
Confidence3
Gert LabsV3 evidence
LMSpeed rank#4
Score64.9
UpdatedSep 2, 2026
Confidence3
ResearchClawBenchV3 evidence
LMSpeed rank#13
Score15.3
UpdatedSep 2, 2026
Confidence3
JobBenchV3 evidence
LMSpeed rank#6
Score38.9
UpdatedSep 2, 2026
Confidence3
ExploitGymV3 evidence
LMSpeed rank#5
Score6.0
UpdatedSep 2, 2026
Confidence3

Coding

V3.0

undefined metric} other undefined metrics}} · Estimated

Score62.280% interval51.9–72.52/4 Measured dimensionsSciCode47.1%#35 / 207Coding score42.9#71 / 82
Dimensions and evidence
Code generation65.2
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
livecodebench · live_code_bench_pro · z 1.29 · q 0.50
scicode · scicode · z 1.08 · q 1.00
Repository engineering59.2
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
react_native_bench · react_native_evals · z 0.98 · q 1.00
swe_pro · swe_pro · z 0.14 · q 1.00
vibecode · vibe_code_bench · z 1.43 · q 1.00
Debugging & testingPrior only
Tool-assisted development & qualityPrior only
SciCodeV3 evidence
LMSpeed rank#35
Score47.1%
UpdatedSep 2, 2026
Confidence4
Coding score
LMSpeed rank#71
Score42.9
UpdatedSep 2, 2026
Confidence3
LiveCodeBench ProV3 evidence
LMSpeed rank#1
Score87.5
UpdatedSep 2, 2026
Confidence3
SWE-bench ProV3 evidence
LMSpeed rank#21
Score57.7
UpdatedSep 2, 2026
Confidence3
React Native EvalsV3 evidence
LMSpeed rank#1
Score85.3
UpdatedSep 2, 2026
Confidence3
Vibe Code BenchV3 evidence
LMSpeed rank#3
Score67.4
UpdatedSep 2, 2026
Confidence3
AA Coding Index
LMSpeed rank#18
Score71.0
UpdatedSep 2, 2026
Confidence3
AA-SciCodeV3 evidence
LMSpeed rank#7
Score56.6
UpdatedSep 2, 2026
Confidence3
EEBench
LMSpeed rank#9
Score38.5
UpdatedAug 20, 2026
Confidence3

Reasoning

V3.0

undefined metric} other undefined metrics}} · Rated

Score55.9#1580% interval47.2–64.63/4 Measured dimensionsGPQA74.8%#118 / 214HLE11.3%#111 / 211
Dimensions and evidence
Abstract logic51.4
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
arc_agi · arc_agi2 · z 0.04 · q 1.00
Scientific & causal reasoning59.6
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
critpt · critpt · z 0.95 · q 1.00
gpqa · gpqa · z 0.66 · q 1.00
hle · hle · z 0.67 · q 1.00
hle · hle_no_tools · z -0.23 · q 1.00
Multi-step constraintsPrior only
Evidence integration & verification56.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
lcr · lcr · z 0.52 · q 1.00
GPQAV3 evidence
LMSpeed rank#118
Score74.8%
UpdatedSep 2, 2026
Confidence4
HLEV3 evidence
LMSpeed rank#111
Score11.3%
UpdatedSep 2, 2026
Confidence4
AA-LCRV3 evidence
LMSpeed rank#16
Score77.7
UpdatedSep 2, 2026
Confidence3
CritPtV3 evidence
LMSpeed rank#9
Score23.4
UpdatedSep 2, 2026
Confidence3
Reasoning score
LMSpeed rank#13
Score69.8
UpdatedSep 2, 2026
Confidence3
ARC-AGI-2V3 evidence
LMSpeed rank#9
Score74.0
UpdatedSep 2, 2026
Confidence3
ARC-AGI-3
LMSpeed rank#8
Score0.2
UpdatedSep 2, 2026
Confidence3

Knowledge

V3.0

undefined metric} other undefined metrics}} · Provisional

Score59.980% interval44.7–75.01/4 Measured dimensionsKnowledge score75.1#23 / 70HLE w/o tools39.8#15 / 22
Dimensions and evidence
Broad knowledgePrior only
Professional knowledge59.9
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
healthbench · health_bench_hard · z 0.74 · q 0.88
medxpert · med_xpert_qa_text · z 0.45 · q 0.50
FactualityPrior only
Retrieval & open-book usePrior only
Knowledge score
LMSpeed rank#23
Score75.1
UpdatedSep 2, 2026
Confidence3
HLE w/o tools
LMSpeed rank#15
Score39.8
UpdatedSep 2, 2026
Confidence3
GPQA-D
LMSpeed rank#8
Score92.8
UpdatedSep 2, 2026
Confidence3
HealthBench HardV3 evidence
LMSpeed rank#1
Score40.1
UpdatedSep 2, 2026
Confidence3
MedXpertQA (Text)V3 evidence
LMSpeed rank#2
Score59.6
UpdatedSep 2, 2026
Confidence3
Artificial Analysis Intelligence Index
LMSpeed rank#18
Score53.1
UpdatedSep 2, 2026
Confidence3
AA-GPQA Diamond
LMSpeed rank#18
Score92.0
UpdatedSep 2, 2026
Confidence3
AA-HLE
LMSpeed rank#12
Score43.7
UpdatedSep 2, 2026
Confidence3
AA-Omniscience Index
LMSpeed rank#24
Score5.8
UpdatedSep 2, 2026
Confidence3
AA-Omniscience Accuracy
LMSpeed rank#13
Score50.8
UpdatedSep 2, 2026
Confidence3
AA-Omniscience Hallucination Rate
LMSpeed rank#13
Score91.7
UpdatedSep 2, 2026
Confidence3
HealthBench Professional
LMSpeed rank#6
Score48.1
UpdatedSep 2, 2026
Confidence3

Math

V3.0

undefined metric} other undefined metrics}} · Provisional

Score57.680% interval41.5–73.61/4 Measured dimensionsMath score65.7#23 / 62FrontierMath v2 (Tiers 1-3)47.6#7 / 47
Dimensions and evidence
Foundational mathPrior only
Competition mathPrior only
Advanced proofs57.6
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
frontiermath · frontier_math_v2_tier4 · z 0.97 · q 1.00
frontiermath · frontier_math_v2_tiers13 · z 0.61 · q 1.00
Applied & tool-assisted mathPrior only
Math score
LMSpeed rank#23
Score65.7
UpdatedSep 2, 2026
Confidence3
FrontierMath v2 (Tiers 1-3)V3 evidence
LMSpeed rank#7
Score47.6
UpdatedSep 2, 2026
Confidence3
FrontierMath v2 (Tier 4)V3 evidence
LMSpeed rank#8
Score27.1
UpdatedSep 2, 2026
Confidence3

Multilingual

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Cross-language understandingPrior only
Multilingual generationPrior only
Reasoning transferPrior only
Low-resource robustnessPrior only

Multimodal

V3.0

undefined metric} other undefined metrics}} · Rated

Score52.9#480% interval46.2–59.54/4 Measured dimensionsMultimodal Grounded score64.8#19 / 44MMMU-Pro81.2#7 / 31
Dimensions and evidence
Perception & OCR46.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
simplevqa · simple_vqa · z -0.43 · q 1.00
Document & spatial understanding50.4
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
charxiv · charxiv · z 0.18 · q 1.00
Visual reasoning56.9
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
erqa · erqa · z -0.01 · q 1.00
medxpert · med_xpert_qa_mm · z 0.33 · q 0.75
mmmu_pro · mmmu_pro · z 0.69 · q 1.00
Video & grounded action57.6
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
screenspot · screen_spot_pro · z 0.73 · q 1.00
Multimodal Grounded score
LMSpeed rank#19
Score64.8
UpdatedSep 2, 2026
Confidence3
MMMU-ProV3 evidence
LMSpeed rank#7
Score81.2
UpdatedSep 2, 2026
Confidence3
OfficeQA Pro
LMSpeed rank#7
Score53.2
UpdatedSep 2, 2026
Confidence3
MMMU-Pro w/ Python
LMSpeed rank#4
Score82.1
UpdatedSep 2, 2026
Confidence3
CharXivV3 evidence
LMSpeed rank#11
Score82.8
UpdatedSep 2, 2026
Confidence3
ERQAV3 evidence
LMSpeed rank#5
Score65.4
UpdatedSep 2, 2026
Confidence3
SimpleVQAV3 evidence
LMSpeed rank#5
Score61.1
UpdatedSep 2, 2026
Confidence3
ScreenSpot ProV3 evidence
LMSpeed rank#2
Score85.4
UpdatedSep 2, 2026
Confidence3
ZeroBench
LMSpeed rank#1
Score41.0
UpdatedSep 2, 2026
Confidence3
MedXpertQA (MM)V3 evidence
LMSpeed rank#3
Score77.1
UpdatedSep 2, 2026
Confidence3
AA-MMMU-Pro
LMSpeed rank#21
Score78.4
UpdatedSep 2, 2026
Confidence3
Design Arena Website
LMSpeed rank#39
Score1238.0
UpdatedSep 2, 2026
Confidence3

Instruction following

V3.0

undefined metric} other undefined metrics}} · Provisional

Score53.980% interval37.9–69.91/4 Measured dimensionsAA-IFBench73.9#22 / 84
Dimensions and evidence
Constraint followingPrior only
Structured outputPrior only
Novel-instruction generalization53.9
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
ifbench · aa_if_bench · z 0.24 · q 1.00
Multi-turn & long instructionsPrior only
AA-IFBenchV3 evidence
LMSpeed rank#22
Score73.9
UpdatedSep 2, 2026
Confidence3

OpenRouter endpoints

7 endpoints

Third-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.

Provider endpointInputOutput1d uptime30m latency30m throughputContext / output
Amazon Bedrock
amazon-bedrock/us-east-1
$2.75/M$16.50/M100%undefined tokens / undefined tokens
Azure
azure/eu
$2.75/M$16.50/M100%undefined tokens / undefined tokens
Azure
azure/us
$2.75/M$16.50/M100%undefined tokens / undefined tokens
OpenAI
openai/fast
$5/M$30/M100%undefined tokens / undefined tokens
Azure
azure
$2.50/M$15/M99.8%undefined tokens / undefined tokens
OpenAI
openai
$2.50/M$15/M99.1%undefined tokens / undefined tokens
OpenAI
openai/flex
$1.25/M$7.50/M97.1%undefined tokens / undefined tokens

Pricing Comparison

Compare GPT-5.4 API pricing across 1000 providers. Prices range from $0.0027/request to $749250.00/M. KFCV50 offers the lowest rate at $0.0027/request. 10 providers offer free API credits or a free tier.

ProviderHealthModel VariantGroupInput ($/M)Output ($/M)Speed (t/s)First tokenAudit
L1
46%
gpt-5.4
default
$75.00/M
$450.00/M
88.2 t/s
6.08 s
L1
100%
gpt-5.4
default
$2.50/M
Cache read$0.250/M
$15.00/M
86.2 t/s
6.12 s
L1
100%
gpt-5.4-openai-compact
default
$2.50/M
Cache read$0.250/M
$15.00/M
L1
99%
gpt-5.4
default
$2.68/M
Cache read$0.268/M
$16.06/M
71.2 t/s
1.63 s
L1
0%
gpt-5.4
default
$2.50/M
Cache read$0.250/M
$15.00/M
68.0 t/s
2.15 s
L1
0%
gpt-5.4-openai-compact
default
$2.50/M
Cache read$0.250/M
$15.00/M
L1
100%
gpt-5.4
default
$0.073/request
-
59.9 t/s
2.48 s
L1
100%
gpt-5.4-openai-compact
default
$0.730/request
-
L1
99%
gpt-5.4
GPT-Entry
-97%$0.075/M
Cache read$0.0075/M
-97%$0.450/M
53.2 t/s
1.11 s
L1
0%
gpt-5.4
OpenAI
-86%$0.342/M
Cache read$0.034/M
-86%$2.05/M
53.1 t/s
5.11 s
L1
100%
gpt-5.4
default
-86%$0.342/M
Cache read$0.034/M
-86%$2.05/M
52.6 t/s
2.36 s
L1
100%
gpt-5.4-openai-compact
default
-86%$0.342/M
Cache read$0.034/M
-86%$2.05/M
L1
100%
gpt-5.4
default
-86%$0.342/M
Cache read$0.034/M
-86%$2.05/M
51.3 t/s
2.92 s
L1
100%
gpt-5.4-openai-compact
default
$10.27/M
$61.64/M
L1
100%
GPT-5.4
default
$10.27/M
-32%$10.27/M
L1
99%
gpt-5.4
Codex
$375.00/M
$2250.00/M
51.2 t/s
0.97 s
L1
99%
gpt-5.4-openai-compact
Codex
$375.00/M
$2250.00/M
L1
99%
gpt-5.4
default
-97%$0.084/M
Cache read$0.0084/M
-97%$0.503/M
51.0 t/s
2.53 s
L1
0%
gpt-5.4
default
$12.50/M
Cache read$1.25/M
$75.00/M
50.9 t/s
6.04 s
L1
0%
gpt-5.4-openai-compact
default
$12.50/M
Cache read$1.25/MCache write$1.25/MCache write 1h$2.00/M
$75.00/M
Showing 20 model IDs of 269.

Alternatives & Similar Models

Frequently Asked Questions

What benchmark data does GPT-5.4 include?
LMSpeed shows GPT-5.4 benchmark context, API price, output speed, first-token latency, and provider data across 1010 providers when those signals are available.
What is the GPT-5.4 API price?
GPT-5.4 has pricing from 1010 providers, ranging from $0.0027/request to $749250.00/M. KFCV50 has the lowest listed price.
What does the GPT-5.4 API pricing table include?
The GPT-5.4 API pricing table compares 1010 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
Which provider has the cheapest GPT-5.4 API pricing?
KFCV50 currently has the lowest listed GPT-5.4 price at $0.0027/request across 1010 providers.
Can I compare GPT-5.4 API price and speed together?
Yes. LMSpeed shows GPT-5.4 API price, output speed, first-token latency, and provider health on the same page so you can compare cost and performance together.
Is GPT-5.4 API free?
Yes, GPT-5.4 free API options are available through 10 providerundefined other undefined} on LMSpeed, including 180txt API, 兔子API, 兔子API, 兔子API, Moyanjdc API. These providers offer free API credits or a free tier with no per-token charges.
Where can I get GPT-5.4 free API access?
LMSpeed currently lists 10 free API providerundefined other undefined} for GPT-5.4: 180txt API, 兔子API, 兔子API, 兔子API, Moyanjdc API. Check each provider row before using it because free tier limits can change.

Also known as

gpt-5.411/gpt-5.415/gpt-5.464/gpt-5.464/gpt-5.4-2026-03-05

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.Standard benchmark data may include BenchLM and other public sources.

Build with LMSpeed Data

Free

Free public API for LLM pricing, benchmarks & provider data

  • Real-time pricing & availability
  • Speed & latency benchmarks
  • 300+ models, 600+ providers
  • 1,000 requests/day
View API Documentation