Mistral Large 3 API Benchmarks, Pricing & Provider Data
Compare Mistral Large 3 with another model
Choose a model to open its comparison page.
The page also shows measured API speed and first-token latency.
Mistral Large 3 is Mistral AI's advanced multilingual frontier model, offering strong reasoning, coding, and instruction-following for production chat and agent workloads.
Category Performance
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
- Coverage
- 6 / 8
- Methodology
- V3.0
#1Coding47.9Provisional1/4 Measured dimensions
#2Reasoning44.5RatedGlobal rank #553/4 Measured dimensions
#3Math43.5Provisional1/4 Measured dimensions
#4Instruction following42.2Provisional1/4 Measured dimensions
#5Knowledge41.6Provisional1/4 Measured dimensions
#6Agents39.2Estimated2/4 Measured dimensions
Specifications
Rankings
Excels at
Falls behind in
Detailed scores
Updated: Aug 31, 2026Speed & latency
undefined metric} other undefined metrics}}
Output speed37.8 tok/s#69 / 76Time to first token2.36 s#59 / 76
Speed & latency
undefined metric} other undefined metrics}}
Pricing
undefined metric} other undefined metrics}}
Input price$0.500/M#82 / 180Output price$1.50/M#62 / 180
Pricing
undefined metric} other undefined metrics}}
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Score39.280% interval27.4–51.12/4 Measured dimensionsAA Agentic Index5.5#56 / 65Τ²-bench results24.6#74 / 82
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Coding
V3.0undefined metric} other undefined metrics}} · Provisional
Score47.980% interval33.9–61.81/4 Measured dimensionsLiveCodeBench46.5%#64 / 115SciCode36.2%#118 / 205
Coding
V3.0undefined metric} other undefined metrics}} · Provisional
Reasoning
V3.0undefined metric} other undefined metrics}} · Rated
Score44.5#5580% interval35.9–53.23/4 Measured dimensionsMMLU-Pro80.7%#58 / 129GPQA68.0%#141 / 212
Reasoning
V3.0undefined metric} other undefined metrics}} · Rated
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Score41.680% interval25.6–57.61/4 Measured dimensionsArtificial Analysis Intelligence Index15.9#92 / 111AA-GPQA Diamond68.0#89 / 108
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Math
V3.0undefined metric} other undefined metrics}} · Provisional
Score43.580% interval27.4–59.61/4 Measured dimensions
Math
V3.0undefined metric} other undefined metrics}} · Provisional
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
80% interval30.8–69.20/4 Measured dimensionsAA-MMMU-Pro55.7#58 / 65
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
Score42.280% interval26.2–58.21/4 Measured dimensionsAA-IFBench36.2#77 / 84
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
Alternatives & Similar Models
GLM-5.1
glm-5-1
Zhipu GLM-5.1 is a next-generation GLM model aimed at frontier reasoning, coding, and bilingual agent applications.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
Gemini 3 Flash
gemini-3-flash
Google Gemini 3 Flash is a next-generation fast multimodal model for responsive assistants, document understanding, and high-throughput API traffic.
MiniMax M2.7
minimax-m2-7
MiniMax M2.7 is a high-tier M2-series model tuned for complex reasoning, long-context dialogue, and production-grade API workloads.
GPT-OSS
gpt-oss
GPT-OSS is an open-weight language model family designed for self-hosted inference, research, and cost-efficient alternatives to proprietary GPT-class models.
DeepSeek V4 Pro
deepseek-v4-pro
DeepSeek V4 Pro is the professional-tier DeepSeek V4 model, targeting frontier reasoning, coding, and agent workflows with maximum capability.
