Category Score V3 leaderboard

LMSpeed Best Models for Math

Compare the best AI models for math across foundational, competition, proof, and tool-assisted mathematics benchmarks, with raw results shown beside the LMSpeed category estimate.

Methodology 3.0Methodology

Current answer

No model currently meets the formal ranking rules. The table has 89 Estimated models and 11 Provisional models. They are useful signals, but they are not formal ranks.

Rankings can change when data or methods change. The run date appears above.

Available leaderboard data

Models shown
100
Formally ranked models
0
Benchmark columns
13
Dimensions with evidence
4/4

How to read the benchmark bars

Each bar compares models only within the same benchmark column. Bar lengths are relative to the models shown here; they are not Category Scores and cannot be compared across benchmark columns.

RankModelLMSpeed scoreFoundational mathCompetition mathAdvanced proofsApplied & tool-assisted mathStatusEvidenceUpdated
MATH-500undefined modelsAA Math Indexundefined modelsAIMEundefined modelsHMMT Feb 2026undefined modelsAIME26undefined modelsHMMT Nov 2025undefined modelsHMMT Feb 2025undefined modelsAIME 2025undefined modelsFrontierMath v2 (Tiers 1-3)undefined modelsFrontierMath v2 (Tier 4)undefined modelsFrontierMath (legacy)undefined modelsIMOAnswerBenchundefined modelsMMAnswerBenchundefined models
Estimated models — unranked89
GLM-5.2Z.ai
70.1±11.5
92.599.294.491.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
O4 MiniOpenAI
63.9±11.8
98.9%94.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3 235B A22B Instruct 2507Qwen
62.9±11.8
98.4%94.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.7 MaxQwen
62.2±12.2
97.190.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5OpenAI
61.4±11.8
98.7%83.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
o3 Mini HighOpenAI
61.3±11.8
98.5%86.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 2.5 Pro Preview 06-05Google
60.7±11.8
98.0%87.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GLM-4.5Z.ai
60.5±11.8
97.9%87.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
DeepSeek V4 Pro 0813DeepSeek
58.9±12.2
95.289.8Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
MiniMax M1MiniMax
58.9±11.8
97.2%81.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
O3OpenAI
58.8±9.3
99.2%90.3%18.72.1Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Kimi K2.6MoonshotAI
58.6±9.0
92.796.439.014.686.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
O3 MiniOpenAI
58.5±11.8
97.3%77.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3
58.3±11.8
97.5%72.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
DeepSeek V4 Flash 0731DeepSeek
57.9±12.2
94.888.4Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Sonar Reasoning ProPerplexity
57.4±11.8
95.7%79.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
DeepSeek R1DeepSeek
57±11.8
96.6%68.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GLM-4.5 AirZ.ai
56.9±11.8
96.5%67.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 2.5 ProGoogle
56.1±9.3
96.7%88.7%14.14.2Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
DeepSeek R1 Distill QwenDeepSeek
55.7±11.8
94.9%66.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.7 PlusQwen
55.2±12.2
92.986.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
R1 Distill Llama 70BDeepSeek
55±11.8
93.5%67.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
O1 MiniOpenAI
54.9±11.8
94.4%60.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4Anthropic
54.4±11.8
94.1%56.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Kimi K2MoonshotAI
53.7±9.3
97.1%69.3%21.40.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 2.5 Flash LiteGoogle
53.3±11.8
92.6%50.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Sonnet 4Anthropic
52.9±11.8
93.4%40.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
O1OpenAI
52.6±9.3
92.4%72.3%9.3Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Reka Flash 3Rekaai
52.2±11.8
89.3%51.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.6 PlusQwen
52.2±9.1
87.895.394.696.726.28.383.8Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Mistral Medium 3Mistral
52.1±11.8
90.7%44.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 2.0 FlashGoogle
52.1±11.8
93.0%33.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 2.0 ProGoogle
52.1±11.8
92.3%36.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GLM-5.1Z.ai
51.7±9.1
82.695.394.033.412.583.8Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GLM-5Z.ai
51.1±9.1
86.495.896.997.516.42.182.5Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3 235B A22BQwen
51.1±11.8
90.2%32.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 2.5 FlashGoogle
50.8±9.3
92.6%43.3%4.84.2Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3 Coder 30B A3B InstructQwen
50.5±11.8
89.3%29.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
SonarPerplexity
50.3±11.8
81.7%48.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 2.0 Flash LiteGoogle
50.1±11.8
87.3%30.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3 32BQwen
49.9±11.8
86.9%30.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-4.1 MiniOpenAI
49.9±9.3
92.5%43.0%4.5Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4.5Anthropic
49.9±9.1
85.395.193.392.920.74.284.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 2.0 Flash Lite 001Google
49.8±11.8
87.3%27.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3 14BQwen
49.8±11.8
87.1%28.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
DeepSeek V3
49.5±9.3
94.2%52.0%1.7Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3 30B A3BQwen
49.4±11.8
86.3%26.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Kimi K2.5MoonshotAI
49.2±9.0
87.195.891.195.496.127.94.281.8Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-4.1OpenAI
49±9.3
91.3%43.7%5.50.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GLM-4.7Z.ai
48.8±12.3
95.72.40.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude 3.7 SonnetAnthropic
48.7±11.8
85.0%22.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3 8BQwen
48.5±11.8
82.8%24.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Sonar ProPerplexity
47.5±11.8
74.5%29.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Grok 3
47.3±9.3
87.0%33.0%3.80.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Phi 4Microsoft
46.9±11.8
81.0%14.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Ling-3.0-flashInclusionai
46.7±11.5
87.093.283.7Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Command ACohere
46.2±11.8
81.9%9.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Grok-2
46.2±11.8
77.8%13.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-4o (2024-08-06)OpenAI
46.2±11.8
79.5%11.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Llama 4 MaverickMeta
46.2±9.3
88.9%39.0%0.7Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-4o MiniOpenAI
46.1±11.8
78.9%11.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-4o (2024-05-13)OpenAI
46±11.8
79.1%11.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-4 TurboOpenAI
45.8±11.8
73.7%15.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
DeepSeek R1 Distill Qwen 1.5BDeepSeek
45.5±11.8
68.7%17.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-4.1 NanoOpenAI
45.2±9.3
84.8%23.7%1.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Mistral SabaMistral
44.7±11.8
67.7%13.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Mistral Large 2407Mistralai
44.5±11.8
71.4%9.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Llama 4 ScoutMeta
44.4±9.3
84.4%28.3%0.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Phi 4 Multimodal Instruct
44.2±11.8
69.3%9.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 1.5 ProGoogle
43.7±11.8
67.3%8.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Devstral SmallMistral AI
43.4±11.8
68.4%6.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude 3.5 SonnetAnthropic
43±9.3
69.5%9.7%2.10.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 1.5 FlashGoogle
42.6±11.8
55.4%9.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude 3.5 HaikuAnthropic
42.5±11.8
72.1%3.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Sonnet 4.5Anthropic
42.5±12.3
87.013.54.2Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-4oOpenAI
42.2±9.3
79.7%10.3%0.3Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Phi 4 Mini Instruct
42±11.8
69.6%3.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude 3 OpusAnthropic
41.5±11.8
64.1%3.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.5
41.4±11.5
87.993.392.794.880.9Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.6 27BQwen
41.3±11.5
84.394.190.793.880.8Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Hermes 3 70B InstructNous
39.6±11.8
53.8%2.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude 3 SonnetAnthropic
39.6±11.8
41.4%4.7%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude 2.1Anthropic
38.5±11.8
37.4%3.3%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Mixtral 8x22B InstructMistral
36.8±11.8
54.5%0.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Mistral LargeMistralai
36.6±11.8
52.7%0.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude 3 HaikuAnthropic
36.4±11.8
39.4%1.0%Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.6 35B A3BQwen
34.6±11.5
83.692.789.190.778.9Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
DeepSeek V4 FlashDeepSeek
32.6±12.2
40.841.9Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
DeepSeek V4 ProDeepSeek
31.6±12.2
31.735.3Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Provisional models — unranked11
GPT-5.6 SolOpenAI
67±16.1
89.083.089.0Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
InklingThinking Machines
65.4±16.3
97.1Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.6 TerraOpenAI
64.5±16.1
84.968.384.9Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.6 LunaOpenAI
62.4±16.1
78.658.578.6Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4.8Anthropic
58±16.1
47.231.3Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.4OpenAI
57.6±16.1
47.627.1Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.5 ProOpenAI
57.2±16.1
51.039.652.4Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.5OpenAI
56.9±16.1
51.735.451.7Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.4 ProOpenAI
56.8±16.1
50.037.550.0Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4.7Anthropic
56.8±16.1
43.822.9Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4.6Anthropic
56.5±16.1
40.722.9Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026

What this leaderboard measures

Which AI model is better suited to math tasks?

This leaderboard covers foundational math, competition problems, advanced proofs, and tool-assisted mathematics. Different math tasks require different capabilities.

Available data covers 4/4 dimensions and shows 13 benchmark columns.

Four capability dimensions

The four dimensions come from the category blueprint. Available data may cover only some of them. A dimension without evidence is not presented as a verified capability.

Foundational math

2 benchmark columns currently provide evidence for this dimension.

  • MATH-500
  • AA Math Index

Competition math

6 benchmark columns currently provide evidence for this dimension.

  • AIME
  • HMMT Feb 2026
  • AIME26
  • HMMT Nov 2025
  • HMMT Feb 2025
  • AIME 2025

Advanced proofs

4 benchmark columns currently provide evidence for this dimension.

  • FrontierMath v2 (Tiers 1-3)
  • FrontierMath v2 (Tier 4)
  • FrontierMath (legacy)
  • IMOAnswerBench

Applied & tool-assisted math

1 benchmark columns currently provide evidence for this dimension.

  • MMAnswerBench

Math tasks this page can help with

  • Foundational problems in arithmetic, algebra, and common school mathematics.
  • Competition problems and proofs that require a valid multi-step solution.
  • Applied calculations where the model must set up formulas and check tool output.

How to choose a model with this leaderboard

  1. Step 1

    Check the rating status first

    Only Rated models receive a rank. Estimated and Provisional models do not have a formal position.

  2. Step 2

    Review uncertainty and evidence

    When scores are close, do not rely on rank alone. Check uncertainty, dimension coverage, and benchmark count.

  3. Step 3

    Test the real task last

    A leaderboard cannot replace your own test. Check quality, speed, price, context, and provider limits together.

Choose by problem type. Recalculate important values and check every proof step.

Rating status guide

Rated

Rated means the evidence and overlap rules are met. The model can receive a formal rank.

Estimated

Estimated means there is useful evidence, but it is not enough for a formal rank.

Provisional

Provisional means evidence is limited or dimension and benchmark-family coverage is below the estimated threshold. Use the result only as an early signal.

Benchmarks and evidence sources

Evidence source names and benchmark groups come from the currently available score data. One source may contribute several benchmarks.

  • Artificial Analysis

    AA Math Index, AIME, and MATH-500

  • BenchLM

    AIME 2025, AIME26, FrontierMath (legacy), FrontierMath v2 (Tier 4), FrontierMath v2 (Tiers 1-3), HMMT Feb 2025, HMMT Feb 2026, HMMT Nov 2025, IMOAnswerBench, and MMAnswerBench

How the math model ranking is built

LMSpeed combines eligible third-party benchmarks inside four fixed capability dimensions. Rated models meet the evidence and overlap requirements for a formal rank; Estimated and Provisional models remain visible without receiving a rank.

Read the Category Score methodology

Leaderboard limits

Category Scores use the third-party benchmarks currently included by LMSpeed. Tests can use different data, prompts, and scoring rules. The result is not permanent and cannot represent every real task. Test important choices with your own data and workflow.

Frequently asked questions

Which visible model has the highest formal rank now?

There is no formal number one now. The page has 89 Estimated models and 11 Provisional models. They do not have a formal rank and should not be called the winner.

Can I compare scores across different categories?

No. Each category uses different capability dimensions and evidence. A Category Score is comparable only inside the same leaderboard. Review the matching category for each task.

Are Estimated and Provisional models still useful?

They can help you find candidates, but their evidence is not complete enough for a formal rank. Review coverage and uncertainty, then test the model on a real task.

How often does the leaderboard update?

The leaderboard updates after a new completed score run is published. The current run date and methodology version appear above. LMSpeed does not promise a fixed daily or weekly schedule.

Is the number one model always best for me?

No. Your result also depends on speed, price, context length, tool support, region, and provider limits. Use the leaderboard to narrow the field, then run your own test.

Does the math ranking represent general reasoning?

No. The math leaderboard describes performance on math tasks. Logic, causality, and evidence verification belong in the reasoning leaderboard. Real work may also require tool use and stable answers.