OpenAI
·Released on Sep 4, 2026

GPT-6 Astra API Benchmarks, Pricing & Provider Data

Compare GPT-6 Astra with another model

Choose a model to open its comparison page.

Share on X
LLM

GPT-6 Astra benchmark, API pricing, and provider data cover 4 API providers, with prices starting at $3.80/M.

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular s...

Quality
#2of 110
81.0
LMSpeed score
Cost
#174of 185
$3.80/ 1M · 8:1 in:out
$0.608 in · $3.19 out

Category Performance

Observed-capability estimates with uncertainty reported separately from independent benchmark families.

Coverage
6 / 8
Methodology
V3.0
Category PerformanceObserved-capability estimates with uncertainty reported separately from independent benchmark families.Agents62.9Coding58.9Reasoning61.5Knowledge58Math72.3Multilingual-Multimodal66.4Instruction following-
#1Math72.3Provisional1/4 Measured dimensions
80% interval: 56.288.5
foundational mathPrior only
competition mathPrior only
proof frontier72.3
frontiermath · frontier_math_v2_tier4
applied tool mathPrior only
#2Multimodal66.4Provisional1/4 Measured dimensions
80% interval: 50.182.8
perception ocrPrior only
document spatialPrior only
visual reasoningPrior only
video action66.4
screenspot · screen_spot_pro
#3Agents62.9RatedGlobal rank #63/4 Measured dimensions
80% interval: 54.671.3
planningPrior only
tool use55
tau · aa_tau3_banking
environment execution67.3
browsecomp · browse_comp / gdpval_aa · benchlm_agentic_gdpval_aa / osworld · os_world2
recovery reliability66.6
briefcase · aa_briefcase_elo / exploit_gym · exploit_gym
#4Reasoning61.5RatedGlobal rank #13/4 Measured dimensions
80% interval: 52.870.2
abstract logic62.9
arc_agi · arc_agi2
scientific causal67
critpt · critpt / gpqa · gpqa / hle · hle
multistep constraintsPrior only
evidence verification54.6
lcr · lcr
#5Coding58.9Estimated2/4 Measured dimensions
80% interval: 47.070.8
code generation59.4
scicode · scicode
repository engineeringPrior only
debugging testingPrior only
tooling quality58.4
terminalbench · aa_terminal_bench21
#6Knowledge58Provisional1/4 Measured dimensions
80% interval: 41.474.6
broad knowledgePrior only
professional knowledge58
healthbench · health_bench_hard
factualityPrior only
retrieval open bookPrior only
No data:MultilingualInstruction following

Specifications

Input and output token limits for this model, plus how it ranks on long-context understanding.

INPUT
1.1Mtokens
1.3K pages of text
OUTPUT
128Ktokens
8K128K1M4M
1.1M

Features

Technical Details

Input
Output
Released
Sep 2026
Documentation
Tokenizer
GPT
Architecture
text+image+file->text
Moderated
Yes
Supported parameters
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Rankings

Excels at

It's decent at

Falls behind in

Detailed scores

Updated: Sep 5, 2026

Overall

undefined metric} other undefined metrics}}

Overall score81.0#2 / 110DeepSWE74.1#2 / 21
Overall score
LMSpeed rank#2
Score81.0
UpdatedSep 5, 2026
Confidence3
DeepSWE
LMSpeed rank#2
Score74.1
UpdatedSep 5, 2026
Confidence3
ExploitBench
LMSpeed rank#1
Score100.0
UpdatedSep 5, 2026
Confidence3
SRE-Bench
LMSpeed rank#1
Score88.0
UpdatedSep 5, 2026
Confidence3
FrontierCyber
LMSpeed rank#1
Score38.1
UpdatedSep 5, 2026
Confidence3
CyScenarioBench success
LMSpeed rank#1
Score59.0
UpdatedSep 5, 2026
Confidence3
CyScenarioBench solved
LMSpeed rank#1
Score9.0
UpdatedSep 5, 2026
Confidence3
Atomic network attacks
LMSpeed rank#1
Score100.0
UpdatedSep 5, 2026
Confidence3
Atomic vulnerability research
LMSpeed rank#1
Score100.0
UpdatedSep 5, 2026
Confidence3
Atomic evasion
LMSpeed rank#3
Score52.0
UpdatedSep 5, 2026
Confidence3
ACCR standard
LMSpeed rank#1
Score3.5
UpdatedSep 5, 2026
Confidence3
ACCR Daybreak Blue
LMSpeed rank#1
Score3.5
UpdatedSep 5, 2026
Confidence3
SEC-Bench Pro
LMSpeed rank#1
Score85.4
UpdatedSep 5, 2026
Confidence3

Pricing

undefined metric} other undefined metrics}}

Input price$10.00/M#174 / 185Output price$50.00/M#175 / 185
Input price
LMSpeed rank#174
Score$10.00/M
UpdatedSep 5, 2026
Confidence4
Output price
LMSpeed rank#175
Score$50.00/M
UpdatedSep 5, 2026
Confidence4

Agents

V3.0

undefined metric} other undefined metrics}} · Rated

Score62.9#680% interval54.6–71.33/4 Measured dimensionsAgentic score87.1#5 / 76OSWorld 2.072.6#1 / 19
Dimensions and evidence
Planning & decompositionPrior only
Tool use55
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
tau · aa_tau3_banking · z 0.16 · q 1.00
Environment & long-horizon execution67.3
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
browsecomp · browse_comp · z 1.24 · q 1.00
gdpval_aa · benchlm_agentic_gdpval_aa · z 1.06 · q 1.00
osworld · os_world2 · z 1.53 · q 1.00
Recovery & completion reliability66.6
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
briefcase · aa_briefcase_elo · z 0.58 · q 1.00
exploit_gym · exploit_gym · z 1.49 · q 0.58
Agentic score
LMSpeed rank#5
Score87.1
UpdatedSep 5, 2026
Confidence3
OSWorld 2.0V3 evidence
LMSpeed rank#1
Score72.6
UpdatedSep 5, 2026
Confidence3
Terminal-Bench 4.0
LMSpeed rank#1
Score57.9
UpdatedSep 5, 2026
Confidence3
Terminal-Bench-Science 0.1
LMSpeed rank#1
Score64.6
UpdatedSep 5, 2026
Confidence3
ExploitGymV3 evidence
LMSpeed rank#1
Score42.4
UpdatedSep 5, 2026
Confidence3
Agents' Last Exam
LMSpeed rank#1
Score59.3
UpdatedSep 5, 2026
Confidence3
AA Agentic Index
LMSpeed rank#9
Score51.5
UpdatedSep 5, 2026
Confidence3
GDPval-AA
LMSpeed rank#9
Score56.5
UpdatedSep 5, 2026
Confidence3
GDPval-AAV3 evidence
LMSpeed rank#11
Score1629.0
UpdatedSep 5, 2026
Confidence3
AA BriefcaseV3 evidence
LMSpeed rank#3
Score1569.0
UpdatedSep 5, 2026
Confidence3
AA Tau3 BankingV3 evidence
LMSpeed rank#10
Score41.4
UpdatedSep 5, 2026
Confidence3
BrowseCompV3 evidence
LMSpeed rank#2
Score91.5
UpdatedSep 5, 2026
Confidence3
AutomationBench
LMSpeed rank#3
Score41.4
UpdatedSep 5, 2026
Confidence3
HLE w/ tools
LMSpeed rank#4
Score57.2
UpdatedSep 5, 2026
Confidence3
Terminal-Bench 2.1 (Vals)
LMSpeed rank#1
Score87.3
UpdatedSep 5, 2026
Confidence3

Coding

V3.0

undefined metric} other undefined metrics}} · Estimated

Score58.980% interval47.0–70.82/4 Measured dimensionsSciCode53.5%#16 / 37AA Terminal-Bench 2.188.4#3 / 22
Dimensions and evidence
Code generation59.4
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
scicode · scicode · z 1.10 · q 1.00
Repository engineeringPrior only
Debugging & testingPrior only
Tool-assisted development & quality58.4
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
terminalbench · aa_terminal_bench21 · z 0.62 · q 1.00
SciCodeV3 evidence
LMSpeed rank#16
Score53.5%
UpdatedSep 5, 2026
Confidence4
AA Terminal-Bench 2.1V3 evidence
LMSpeed rank#3
Score88.4
UpdatedSep 5, 2026
Confidence3
AA Coding Index
LMSpeed rank#4
Score76.9
UpdatedSep 5, 2026
Confidence3
AA-SciCodeV3 evidence
LMSpeed rank#16
Score54.1
UpdatedSep 5, 2026
Confidence3
Coding score
LMSpeed rank#4
Score78.5
UpdatedSep 5, 2026
Confidence3
FrontierCode 1.1 Main
LMSpeed rank#3
Score53.3
UpdatedSep 5, 2026
Confidence3
FrontierCode 1.1 Extended
LMSpeed rank#1
Score64.5
UpdatedSep 5, 2026
Confidence3

Reasoning

V3.0

undefined metric} other undefined metrics}} · Rated

Score61.5#180% interval52.8–70.23/4 Measured dimensionsGPQA89.5%#40 / 217HLE37.1%#36 / 214
Dimensions and evidence
Abstract logic62.9
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
arc_agi · arc_agi2 · z 1.63 · q 1.00
Scientific & causal reasoning67
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
critpt · critpt · z 1.03 · q 1.00
gpqa · gpqa · z 1.87 · q 1.00
hle · hle · z 1.28 · q 1.00
Multi-step constraintsPrior only
Evidence integration & verification54.6
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
lcr · lcr · z 0.18 · q 1.00
GPQAV3 evidence
LMSpeed rank#40
Score89.5%
UpdatedSep 5, 2026
Confidence4
HLEV3 evidence
LMSpeed rank#36
Score37.1%
UpdatedSep 5, 2026
Confidence4
Reasoning score
LMSpeed rank#1
Score88.8
UpdatedSep 5, 2026
Confidence3
ARC-AGI-2V3 evidence
LMSpeed rank#1
Score95.0
UpdatedSep 5, 2026
Confidence3
ARC-AGI-3
LMSpeed rank#1
Score62.7
UpdatedSep 5, 2026
Confidence3
MRCR v2 256K-512K
LMSpeed rank#1
Score100.0
UpdatedSep 5, 2026
Confidence3
MRCR v2 512K-1M
LMSpeed rank#2
Score96.3
UpdatedSep 5, 2026
Confidence3
AA-LCRV3 evidence
LMSpeed rank#38
Score74.3
UpdatedSep 5, 2026
Confidence3
CritPtV3 evidence
LMSpeed rank#2
Score31.7
UpdatedSep 5, 2026
Confidence3
ARC-AGI-1
LMSpeed rank#1
Score98.5
UpdatedSep 5, 2026
Confidence3
GeneBench-Pro
LMSpeed rank#1
Score37.8
UpdatedSep 5, 2026
Confidence3

Knowledge

V3.0

undefined metric} other undefined metrics}} · Provisional

Score5880% interval41.4–74.61/4 Measured dimensionsKnowledge score86.8#5 / 82GPQA-D96.0#1 / 32
Dimensions and evidence
Broad knowledgePrior only
Professional knowledge58
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
healthbench · health_bench_hard · z 0.35 · q 1.00
FactualityPrior only
Retrieval & open-book usePrior only
Knowledge score
LMSpeed rank#5
Score86.8
UpdatedSep 5, 2026
Confidence3
GPQA-D
LMSpeed rank#1
Score96.0
UpdatedSep 5, 2026
Confidence3
HealthBench Professional
LMSpeed rank#1
Score63.4
UpdatedSep 5, 2026
Confidence3
HealthBench HardV3 evidence
LMSpeed rank#2
Score36.3
UpdatedSep 5, 2026
Confidence3
Artificial Analysis Intelligence Index
LMSpeed rank#5
Score61.2
UpdatedSep 5, 2026
Confidence3
AA-GPQA Diamond
LMSpeed rank#1
Score96.1
UpdatedSep 5, 2026
Confidence3
AA-HLE
LMSpeed rank#4
Score54.7
UpdatedSep 5, 2026
Confidence3
AA-Omniscience Index
LMSpeed rank#2
Score43.4
UpdatedSep 5, 2026
Confidence3
AA-Omniscience Accuracy
LMSpeed rank#3
Score62.6
UpdatedSep 5, 2026
Confidence3
AA-Omniscience Hallucination Rate
LMSpeed rank#75
Score51.3
UpdatedSep 5, 2026
Confidence3

Math

V3.0

undefined metric} other undefined metrics}} · Provisional

Score72.380% interval56.2–88.51/4 Measured dimensionsMath score85.2#4 / 63FrontierMath v2 (Tier 4)97.6#1 / 36
Dimensions and evidence
Foundational mathPrior only
Competition mathPrior only
Advanced proofs72.3
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
frontiermath · frontier_math_v2_tier4 · z 3.00 · q 1.00
Applied & tool-assisted mathPrior only
Math score
LMSpeed rank#4
Score85.2
UpdatedSep 5, 2026
Confidence3
FrontierMath v2 (Tier 4)V3 evidence
LMSpeed rank#1
Score97.6
UpdatedSep 5, 2026
Confidence3

Multilingual

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Cross-language understandingPrior only
Multilingual generationPrior only
Reasoning transferPrior only
Low-resource robustnessPrior only

Multimodal

V3.0

undefined metric} other undefined metrics}} · Provisional

Score66.480% interval50.1–82.81/4 Measured dimensionsScreenSpot Pro92.7#1 / 12AA-MMMU-Pro86.9#1 / 68
Dimensions and evidence
Perception & OCRPrior only
Document & spatial understandingPrior only
Visual reasoningPrior only
Video & grounded action66.4
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
screenspot · screen_spot_pro · z 1.85 · q 1.00
ScreenSpot ProV3 evidence
LMSpeed rank#1
Score92.7
UpdatedSep 5, 2026
Confidence3
AA-MMMU-Pro
LMSpeed rank#1
Score86.9
UpdatedSep 5, 2026
Confidence3
Multimodal Grounded score
LMSpeed rank#8
Score82.6
UpdatedSep 5, 2026
Confidence3
BenchCAD Vision2Code (tools)
LMSpeed rank#1
Score1.0
UpdatedSep 5, 2026
Confidence3

Instruction following

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Constraint followingPrior only
Structured outputPrior only
Novel-instruction generalizationPrior only
Multi-turn & long instructionsPrior only

Pricing Comparison

Compare GPT-6 Astra API pricing across 4 providers. Prices range from $3.80/M to $9.80/M. 无限智能 offers the lowest rate at $3.80/M.

ProviderHealthModel VariantGroupInput ($/M)Output ($/M)Speed (t/s)First tokenAudit
L1
100%
gpt-6-astra
GPT-PLUS稳定🐢-x0.38-对话 · ≤272K
-62%$3.80/M
Cache read$0.380/MCache write$4.75/M
-62%$19.00/M

Alternatives & Similar Models

Frequently Asked Questions

What benchmark data does GPT-6 Astra include?
LMSpeed shows GPT-6 Astra benchmark context, API price, output speed, first-token latency, and provider data across 4 providers when those signals are available.
What is the GPT-6 Astra API price?
GPT-6 Astra has pricing from 4 providers, ranging from $3.80/M to $9.80/M. 无限智能 has the lowest listed price.
What does the GPT-6 Astra API pricing table include?
The GPT-6 Astra API pricing table compares 4 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
Which provider has the cheapest GPT-6 Astra API pricing?
无限智能 currently has the lowest listed GPT-6 Astra price at $3.80/M across 4 providers.
Is GPT-6 Astra API free?
GPT-6 Astra does not currently have a free API tier on LMSpeed. All 4 providers charge per token.

Also known as

gpt-6-astraopenai/gpt-6-astraopenai/gpt-6-astra:batch

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.Standard benchmark data may include BenchLM and other public sources.

Build with LMSpeed Data

Free

Free public API for LLM pricing, benchmarks & provider data

  • Real-time pricing & availability
  • Speed & latency benchmarks
  • 300+ models, 600+ providers
  • 1,000 requests/day
View API Documentation