Sponsored byFusecodeEnterprise coding API for Claude Code, Codex, and model workflows.
LogoLMSpeed
  • Free
  • Models
  • Providers
  • Leaderboard
LogoLMSpeed
  1. Home
  2. Leaderboard
  3. Best Model For Agent
LogoLMSpeed

The best API speed test tool

GitHubGitHubTwitterX (Twitter)Email
Product
  • Features
  • Pricing
  • FAQ
Leaderboard
  • Overview
  • Speed Ranking
  • Latency Ranking
  • Health Ranking
  • Model Pricing
  • Model Speed
  • Reasoning
  • Coding
Models
  • All Models
  • GPT
  • Claude
  • Gemini
  • DeepSeek
  • Llama
  • Qwen
Free Models
  • All Free Models
  • Free GPT
  • Free Claude
  • Free Gemini
  • Free DeepSeek
  • Free Llama
  • Free Qwen
Tools
  • Speed Test
  • Provider Audit
Company
  • About
Resources
  • Provider Directory
  • Documentation
  • Public API
  • Botab
  • VidBee
Legal
  • Cookie Policy
  • Privacy Policy
  • Terms of Service
© 2026 LMSpeed All Rights Reserved.Made by Nexmoe with ❤️
AgentsCodingReasoningKnowledgeMathMultilingualMultimodalInstruction following

Category Score V3 leaderboard

LMSpeed Best Models for AI Agents

Compare the best AI models for agents across planning, tool use, environment execution, and recovery benchmarks. Formal ranks use LMSpeed Category Score V3 evidence and uncertainty rules.

Updated August 9, 2026·Methodology 3.0·Methodology

Current answer

Among the currently visible formally ranked models, Kimi K3 has the highest position at global rank 1. Its Category Score is 68.9, with an 80% uncertainty range of ±7.3. 55 visible models have a formal rank. This result applies only to this score run.

View Kimi K3 detailsRankings can change when data or methods change. The run date appears above.

Available leaderboard data

Models shown
100
Formally ranked models
55
Benchmark columns
26
Dimensions with evidence
4/4

How to read the benchmark bars

Each bar compares models only within the same benchmark column. Bar lengths are relative to the models shown here; they are not Category Scores and cannot be compared across benchmark columns.

RankModelLMSpeed scorePlanning & decompositionTool useEnvironment & long-horizon executionRecovery & completion reliabilityStatusEvidenceUpdated
Gert Labs50 modelsDeepPlanning1 modelsτ²-bench results84 modelsMCP Atlas28 modelsToolathlon22 modelsAA Tau3 Banking17 modelsτ³-bench results9 modelsMCP-Tasks5 modelsGDPval-AA63 modelsTerminal-Bench 2.051 modelsBrowseComp29 modelsOSWorld-Verified25 modelsAA EnterpriseOps-Gym17 modelsOSWorld 2.014 modelsAA ITBench11 modelsDeepSearchQA11 modelsVITA-Bench9 modelsWideResearch9 modelsClaw-Eval28 modelsAPEX-Agents-AA24 modelsJobBench22 modelsResearchClawBench19 modelsAA Briefcase17 modelsAA Harvey LAB16 modelsCyberGym13 modelsExploitGym6 models
Formally ranked models55
1MoonshotAIKimi K3MoonshotAI
68.9±7.3
———84.2—33.4——1687.088.391.2—45.3—47.795.0———41.352.9—1542.094.6——Rated3/4 dimensions · 12 familiesAug 9, 2026
2OpenAIGPT-5.6 SolOpenAI
68.2±7.3
——85.1—58.033.0——1735.091.992.2—42.962.656.2———————1504.087.284.533.7Rated3/4 dimensions · 12 familiesAug 9, 2026
3ClaudeClaude Opus 5Anthropic
67.5±8.2
———85.8—30.3——1862.0—90.8—47.570.6—95.0——————1720.0———Rated3/4 dimensions · 8 familiesAug 9, 2026
4ClaudeClaude Fable 5Anthropic
65.7±8.2
——98.5——26.8——1747.084.3—85.051.1—————————1574.093.6——Rated3/4 dimensions · 7 familiesAug 9, 2026
5OpenAIGPT-5.6 TerraOpenAI
63.9±7.4
——86.3—53.131.8——1583.087.487.5——50.251.0————38.9———85.281.823.2Rated3/4 dimensions · 11 familiesAug 9, 2026
6ClaudeClaude Opus 4.8Anthropic
63.2±5.3
73.0—94.482.259.927.6——1593.074.684.383.444.020.6—93.1—————21.11345.091.1——Rated4/4 dimensions · 13 familiesAug 9, 2026
7OpenAIGPT-5.5OpenAI
62.7±5.1
72.9—98.075.355.6———1490.082.084.478.746.613.045.8————37.742.717.0——81.813.4Rated4/4 dimensions · 15 familiesAug 9, 2026
8GrokGrok 4.5SpaceXAI
61.9±8.4
—————32.6——1527.083.3——40.8—————————1315.092.4——Rated3/4 dimensions · 6 familiesAug 9, 2026
9GeminiGemini 3.5 FlashGoogle
61.1±5.6
61.9—95.383.656.5———1345.076.2—78.450.1——————47.1—18.0————Rated4/4 dimensions · 10 familiesAug 9, 2026
10ChatGLMGLM-5.2Z.ai
60.5±7.2
——99.176.848.226.8——1510.081.0——42.7—42.7————33.7—20.71253.091.0——Rated3/4 dimensions · 11 familiesAug 9, 2026
11SparkMuse Spark 1.1Meta
60.4±7.1
———88.175.625.2——1375.080.0—80.847.214.2—84.9————54.7—868.093.159.00.8Rated3/4 dimensions · 13 familiesAug 9, 2026
12OpenAIGPT-5.6 LunaOpenAI
59.8±7.4
————53.427.2——1582.084.783.3——45.640.3————35.8———87.977.912.4Rated3/4 dimensions · 11 familiesAug 9, 2026
13ClaudeClaude Sonnet 5Anthropic
59.8±8.2
—————28.2——1603.080.484.781.244.7—————————1385.090.1——Rated3/4 dimensions · 8 familiesAug 9, 2026
14ClaudeClaude Opus 4.7 MaxAnthropic
58±7.7
——88.677.3————1491.069.479.378.0—18.246.7—————45.9———73.1—Rated3/4 dimensions · 9 familiesAug 9, 2026
15OpenAIGPT-5.4OpenAI
57.7±5.1
64.9—98.970.654.6———1391.075.182.775.0———73.6——60.333.338.915.3——79.06.0Rated4/4 dimensions · 15 familiesAug 9, 2026
16QwenQwen3.7 MaxQwen
56.9±5.4
64.3—94.776.4————1270.069.7——45.0—42.5—47.9—65.2——18.7914.083.4——Rated4/4 dimensions · 12 familiesAug 9, 2026
17MinimaxMiniMax M3MiniMax
56.2±7.3
——88.974.2————1390.066.083.570.132.14.6————74.5——19.81108.088.4——Rated3/4 dimensions · 11 familiesAug 9, 2026
18QwenQwen3.6 27BQwen
56.1±6.6
54.8—94.2—————1138.059.3————————72.4———————Rated4/4 dimensions · 5 familiesAug 9, 2026
19MoonshotAIKimi K2.6MoonshotAI
55.5±5.3
56.8—95.955.950.0———1188.066.783.273.1—4.6—92.5—80.862.328.5—18.0————Rated4/4 dimensions · 13 familiesAug 9, 2026
20ClaudeClaude Opus 4.6Anthropic
55.4±5.7
61.9—84.8——————65.483.772.7———73.7——70.433.036.719.9——66.6—Rated4/4 dimensions · 11 familiesAug 9, 2026
21ClaudeClaude Opus 4.7Anthropic
54.9±6.9
65.6—74.0——————————13.9———————20.7————Rated4/4 dimensions · 4 familiesAug 9, 2026
22StepfunStep 3.7 FlashStepFun
54.1±5.7
51.6—98.5—49.5———1017.059.575.8————92.8——67.114.8——————Rated4/4 dimensions · 9 familiesAug 9, 2026
23QwenQwen3.7 PlusQwen
54±5.7
—62.393.073.2————943.070.3—73.3—2.8——45.6—62.722.4——————Rated4/4 dimensions · 9 familiesAug 9, 2026
24ChatGLMGLM-5.1Z.ai
53.9±5.7
60.1—97.771.8——70.6—1256.063.568.0———————62.3——18.2——68.7—Rated4/4 dimensions · 9 familiesAug 9, 2026
25ClaudeClaude Sonnet 4.6Anthropic
53.2±6.1
62.9—79.5——————59.1—72.1—8.3————67.8—36.9———65.2—Rated4/4 dimensions · 7 familiesAug 9, 2026
26GeminiGemini 3.6 FlashGoogle
53±9.0
—————24.5——1423.0——83.0——————————962.0———Rated3/4 dimensions · 4 familiesAug 9, 2026
27OpenAIGPT-5.3 CodexOpenAI
53±6.6
57.5—86.0——————77.3—64.7————————33.7—————Rated4/4 dimensions · 5 familiesAug 9, 2026
28OpenAIGPT-5.4 MiniOpenAI
51.8±8.1
——93.457.742.9———1170.060.0—72.1———————28.2——————Rated3/4 dimensions · 7 familiesAug 9, 2026
29DeepSeekDeepSeek V4 ProDeepSeek
51.4±5.1
50.3—94.269.446.325.8——1293.059.180.4—40.4—38.3———59.824.3—17.1931.084.4——Rated4/4 dimensions · 14 familiesAug 9, 2026
30MiMo-V2.5-ProXiaomi
51.1±5.9
62.7—94.2———72.9—1265.068.4————38.2———63.82.4——879.073.3——Rated4/4 dimensions · 9 familiesAug 9, 2026
31MiMo-V2.5Xiaomi
50.8±8.9
46.9————————65.8————————62.3——16.9————Rated3/4 dimensions · 4 familiesAug 9, 2026
32GeminiGemini 3.1 ProGoogle
50.6±6.2
56.9—95.6—————965.0——————69.7——57.832.0—13.3————Rated4/4 dimensions · 7 familiesAug 9, 2026
33DeepSeekDeepSeek V4 FlashDeepSeek
50.3±5.7
54.4——64.040.731.1——1189.049.153.5———————57.8—————76.7—Rated4/4 dimensions · 9 familiesAug 9, 2026
34QwenQwen3.6 PlusQwen
50±5.5
50.6—97.748.239.8—70.774.11138.061.6——————44.374.358.8——18.0————Rated4/4 dimensions · 11 familiesAug 9, 2026
35InklingThinking Machines
49.4±8.2
———74.1—23.7——1238.063.877.1—38.1—————————840.0———Rated3/4 dimensions · 7 familiesAug 9, 2026
36ClaudeClaude Opus 4.5Anthropic
49.3±5.3
64.2—86.342.343.5—70.271.8—59.3—66.3————23.376.459.6—32.3———50.6—Rated4/4 dimensions · 12 familiesAug 9, 2026
37GrokGrok 4.3SpaceXAI
49.1±6.6
43.9—97.7—————1084.0——————————17.0—12.4————Rated4/4 dimensions · 5 familiesAug 9, 2026
38MiMo-V2-Pro
48±8.9
36.7—95.0———————————————57.8——15.3————Rated3/4 dimensions · 4 familiesAug 9, 2026
39OpenAIGPT-5.2OpenAI
47.9±6.6
46.5—84.8———————65.847.3————————34.3—————Rated4/4 dimensions · 5 familiesAug 9, 2026
40ChatGLMGLM-4.7Z.ai
47.1±8.6
40.0—95.9—————1165.041.052.0—————15.5—————————Rated3/4 dimensions · 6 familiesAug 9, 2026
41ClaudeClaude Sonnet 4.5Anthropic
46.9±8.7
48.5————————50.0—61.4————17.0———27.7—————Rated3/4 dimensions · 5 familiesAug 9, 2026
42QwenQwen3.6 35B A3BQwen
46.8±5.9
42.6—95.362.826.9—67.2—1053.051.5——————35.660.168.7———————Rated4/4 dimensions · 9 familiesAug 9, 2026
43QwenQwen3.5-27BQwen
46.2±8.7
39.4—93.9——————41.661.056.2——————————————Rated3/4 dimensions · 5 familiesAug 9, 2026
44MistralMistral Medium 3.5Mistral
45.9±6.3
39.1—94.2———91.4—932.0———33.7—————————516.069.1——Rated4/4 dimensions · 6 familiesAug 9, 2026
45OpenAIGPT-5.4 NanoOpenAI
45.5±8.1
——92.556.135.5———1101.046.3—39.0———————24.9——————Rated3/4 dimensions · 7 familiesAug 9, 2026
46QwenQwen3.5
45.4±5.3
46.8—95.646.136.3—68.474.2963.052.562.0—————43.774.056.815.3—14.2————Rated4/4 dimensions · 13 familiesAug 9, 2026
47MinimaxMiniMax M2.7MiniMax
45.2±6.0
40.4—84.8—46.3———1158.057.0————————48.710.6——————Rated4/4 dimensions · 7 familiesAug 9, 2026
48GeminiGemini 3 FlashGoogle
43.1±8.9
56.6—43.3———————————————49.2—11.4—————Rated3/4 dimensions · 4 familiesAug 9, 2026
49QwenQwen3.5-35B-A3BQwen
43.1±8.7
29.0—89.2——————40.561.054.5——————————————Rated3/4 dimensions · 5 familiesAug 9, 2026
50GeminiGemini 3.1 Flash LiteGoogle
42.3±6.9
38.5—31.3—————647.0——————————12.2——————Rated4/4 dimensions · 4 familiesAug 9, 2026
51ChatGLMGLM-5Z.ai
41.7±5.6
51.0—98.231.138.0—65.660.8—56.2———————69.857.714.5————43.2—Rated4/4 dimensions · 10 familiesAug 9, 2026
52GeminiGemini 3.5 Flash-LiteGoogle
41.3±8.8
—————16.5——1139.054.0—74.0——————————635.0———Rated3/4 dimensions · 5 familiesAug 9, 2026
53MoonshotAIKimi K2.5MoonshotAI
39.7±5.1
45.9—95.929.527.8—65.759.11003.050.860.6————77.1—72.752.311.58.714.0————Rated4/4 dimensions · 14 familiesAug 9, 2026
54DeepSeekDeepSeek V3.2DeepSeek
39.2±6.9
29.6—78.9—————————————18.5—40.2———————Rated4/4 dimensions · 4 familiesAug 9, 2026
55OpenAIgpt-oss-120bOpenAI
35.3±6.2
29.6—65.8—————802.0———25.5—5.6————3.1———13.9——Rated4/4 dimensions · 7 familiesAug 9, 2026
Estimated models — unranked31
—QwenQwen3.8 MaxQwen
65.8±11.3
———————————86.1—19.4———81.9——53.4—————Estimated2/4 dimensions · 3 familiesAug 9, 2026
—Inkling SmallThinking Machines
55.8±10.9
———79.6————1268.064.777.4———————————————Estimated2/4 dimensions · 4 familiesAug 9, 2026
—MoonshotAIKimi K2.7 CodeMoonshotAI
54±11.2
——90.176.0————1189.0—————————————————Estimated2/4 dimensions · 3 familiesAug 9, 2026
—QwenQwen3.6 Max PreviewQwen
54±11.8
——95.9——————65.4————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—ChatGLMGLM-5 TurboZ.ai
52±11.9
——98.5———————————————55.8———————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—OpenAIGPT-5.2 CodexOpenAI
51.4±9.3
51.8—92.1—————————————————26.0—————Estimated3/4 dimensions · 3 familiesAug 9, 2026
—Ling-3.0-flashInclusionai
50.5±10.2
———65.5—28.0——1107.0—72.2——————73.6————————Estimated2/4 dimensions · 5 familiesAug 9, 2026
—OpenAIGPT-5.1 CodexOpenAI
49.9±9.3
49.7—83.0—————————————————26.2—————Estimated3/4 dimensions · 3 familiesAug 9, 2026
—GeminiGemini 3 ProGoogle
49.4±9.3
63.2—87.1—————————————————11.4—————Estimated3/4 dimensions · 3 familiesAug 9, 2026
—QwenQwen3.5-122B-A10BQwen
48.5±10.6
——93.6—————982.049.463.858.0——————————————Estimated2/4 dimensions · 5 familiesAug 9, 2026
—OpenAIGPT-5.1OpenAI
47.8±9.3
41.2—81.9—————988.0—————————————————Estimated3/4 dimensions · 3 familiesAug 9, 2026
—QwenQwen3 MaxQwen
47.6±11.8
43.7—74.3———————————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—MiMo-V2-Flash
47.3±11.8
——83.9—————838.0—————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—ChatGLMGLM-5V TurboZ.ai
47.1±9.3
30.8—98.5———————————————53.8———————Estimated3/4 dimensions · 3 familiesAug 9, 2026
—HunyuanHy3 previewTencent
46.5±11.2
36.9———————1215.054.4————————————————Estimated2/4 dimensions · 3 familiesAug 9, 2026
—CohereCommand ACohere
46.1±11.8
——85.0—————718.0—————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—OpenAIGPT-5OpenAI
45.6±9.3
——86.5—————1082.0———————————8.5—————Estimated3/4 dimensions · 3 familiesAug 9, 2026
—ClaudeClaude Sonnet 4Anthropic
44.5±9.3
39.7—52.3—————————————————18.4—————Estimated3/4 dimensions · 3 familiesAug 9, 2026
—Ling 2.6 FlashinclusionAI
44.4±11.8
——86.0—————550.0—————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—Trinity Large ThinkingArcee AI
43.9±9.3
32.5—90.1—————564.0—————————————————Estimated3/4 dimensions · 3 familiesAug 9, 2026
—GeminiGemini 2.5 ProGoogle
43.7±9.3
42.0—54.1—————669.0—————————————————Estimated3/4 dimensions · 3 familiesAug 9, 2026
—GrokGrok 4.20SpaceXAI
43.3±11.2
38.4————————47.1—————62.8——————————Estimated2/4 dimensions · 3 familiesAug 9, 2026
—MiMo-V2-Omni
41.5±11.9
——91.2———————————————45.2———————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—OpenAIGPT-4.1 MiniOpenAI
40.4±11.8
——52.9—————505.0—————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—OpenAIGPT-4.1OpenAI
39.9±11.8
25.6—47.1———————————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—MistralMistral Large 3
39.4±11.8
——24.6—————640.0—————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—OpenAIgpt-oss-20bOpenAI
37.7±9.3
——60.2—————564.0——————————0.7——————Estimated3/4 dimensions · 3 familiesAug 9, 2026
—DeepSeekDeepSeek V3
34.6±11.8
——22.8—————231.0—————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—MetaAILlama 4 MaverickMeta
33.3±11.8
——17.8—————5.0—————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—OpenAIGPT-4.1 NanoOpenAI
33.2±11.8
——17.3—————62.0—————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
—MetaAILlama 4 ScoutMeta
32.9±11.8
——15.5—————111.0—————————————————Estimated2/4 dimensions · 2 familiesAug 9, 2026
Provisional models — unranked14
—SparkMuse Spark 1.2Meta
62.3±16.0
————————1631.0—————————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—OpenAIGPT-5.5 ProOpenAI
61.4±16.1
——————————90.1———————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—OpenAIGPT-5.4 ProOpenAI
60.5±16.1
——————————89.3———————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—Laguna S 2.1Poolside
53.7±16.0
—————————70.2————————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—HunyuanHy3Tencent
52.8±16.0
————————1215.0—————————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—OpenAIGPT-5.1 Codex MaxOpenAI
50.1±16.0
——83.0———————————————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—GrokGrok Build 0 1SpaceXAI
50±16.1
49.1—————————————————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—OpenAIO3OpenAI
49.5±16.0
——80.7———————————————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—ChatGLMGLM-4.6Z.ai
48.6±16.0
——76.9———————————————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—ClaudeClaude Opus 4.1Anthropic
46.8±16.2
————————————————————21.9—————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—OpenAIO1OpenAI
45.8±16.0
——62.6———————————————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—MoonshotAIKimi K2MoonshotAI
45.5±16.0
——61.1———————————————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—QwenQwen3.5 PlusAlibaba
44.7±16.2
————————————————————18.5—————Provisional1/4 dimensions · 1 familiesAug 9, 2026
—ChatGLMGLM-4.5 AirZ.ai
43.1±16.0
——46.5———————————————————————Provisional1/4 dimensions · 1 familiesAug 9, 2026

What this leaderboard measures

Which AI model is better suited to agent tasks?

This leaderboard looks at whether a model can plan, use tools, act in an environment, and recover from failure. It measures more than the ability to answer a question.

Available data covers 4/4 dimensions and shows 26 benchmark columns.

Four capability dimensions

The four dimensions come from the category blueprint. Available data may cover only some of them. A dimension without evidence is not presented as a verified capability.

Planning & decomposition

2 benchmark columns currently provide evidence for this dimension.

  • Gert Labs
  • DeepPlanning

Tool use

6 benchmark columns currently provide evidence for this dimension.

  • τ²-bench results
  • MCP Atlas
  • Toolathlon
  • AA Tau3 Banking
  • τ³-bench results
  • MCP-Tasks

Environment & long-horizon execution

10 benchmark columns currently provide evidence for this dimension.

  • GDPval-AA
  • Terminal-Bench 2.0
  • BrowseComp
  • OSWorld-Verified
  • AA EnterpriseOps-Gym
  • OSWorld 2.0
  • AA ITBench
  • DeepSearchQA
  • VITA-Bench
  • WideResearch

Recovery & completion reliability

8 benchmark columns currently provide evidence for this dimension.

  • Claw-Eval
  • APEX-Agents-AA
  • JobBench
  • ResearchClawBench
  • AA Briefcase
  • AA Harvey LAB
  • CyberGym
  • ExploitGym

Agent tasks this page can help with

  • Research assistants that search, organize evidence, and adjust a plan when new facts appear.
  • Browser and computer tasks that require several tool calls and interface actions.
  • Multi-step business automation that must follow rules, handle failure, and leave reviewable output.

How to choose a model with this leaderboard

  1. Step 1

    Check the rating status first

    Only Rated models receive a rank. Estimated and Provisional models do not have a formal position.

  2. Step 2

    Review uncertainty and evidence

    When scores are close, do not rely on rank alone. Check uncertainty, dimension coverage, and benchmark count.

  3. Step 3

    Test the real task last

    A leaderboard cannot replace your own test. Check quality, speed, price, context, and provider limits together.

Start with your tools and execution environment. Then compare recovery, speed, cost, and permission controls. The top-ranked model is not best for every agent.

Rating status guide

Rated

Rated means the evidence and overlap rules are met. The model can receive a formal rank.

Estimated

Estimated means there is useful evidence, but it is not enough for a formal rank.

Provisional

Provisional means evidence is limited or dimension and benchmark-family coverage is below the estimated threshold. Use the result only as an early signal.

Benchmarks and evidence sources

Evidence source names and benchmark groups come from the currently available score data. One source may contribute several benchmarks.

  • BenchLM

    AA Briefcase, AA EnterpriseOps-Gym, AA Harvey LAB, AA ITBench, AA Tau3 Banking, APEX-Agents-AA, BrowseComp, Claw-Eval, CyberGym, DeepPlanning, DeepSearchQA, ExploitGym, GDPval-AA, Gert Labs, JobBench, MCP Atlas, MCP-Tasks, OSWorld 2.0, OSWorld-Verified, ResearchClawBench, Terminal-Bench 2.0, Toolathlon, VITA-Bench, WideResearch, τ²-bench results, and τ³-bench results

How the AI agent model ranking is built

LMSpeed combines eligible third-party benchmarks inside four fixed capability dimensions. Rated models meet the evidence and overlap requirements for a formal rank; Estimated and Provisional models remain visible without receiving a rank.

Read the Category Score methodology

Leaderboard limits

Category Scores use the third-party benchmarks currently included by LMSpeed. Tests can use different data, prompts, and scoring rules. The result is not permanent and cannot represent every real task. Test important choices with your own data and workflow.

Frequently asked questions

Which visible model has the highest formal rank now?

Among the currently visible models, Kimi K3 has the highest formal position at global rank 1. Its Category Score is 68.9. 55 visible models meet the formal ranking rules. This result applies only to the run date and methodology version shown on the page.

Can I compare scores across different categories?

No. Each category uses different capability dimensions and evidence. A Category Score is comparable only inside the same leaderboard. Review the matching category for each task.

Are Estimated and Provisional models still useful?

They can help you find candidates, but their evidence is not complete enough for a formal rank. Review coverage and uncertainty, then test the model on a real task.

How often does the leaderboard update?

The leaderboard updates after a new completed score run is published. The current run date and methodology version appear above. LMSpeed does not promise a fixed daily or weekly schedule.

Is the number one model always best for me?

No. Your result also depends on speed, price, context length, tool support, region, and provider limits. Use the leaderboard to narrow the field, then run your own test.

How is the agent ranking different from the reasoning ranking?

The reasoning ranking focuses on logic, causality, and evidence checks. The agent ranking also covers planning, tool use, environment execution, and recovery. Strong reasoning alone does not prove that a model can finish an agent workflow.