Category Score V3 leaderboard
LMSpeed Best Models for AI Agents
Compare the best AI models for agents across planning, tool use, environment execution, and recovery benchmarks. Formal ranks use LMSpeed Category Score V3 evidence and uncertainty rules.
Methodology 3.0Methodology
Current answer
Among the currently visible formally ranked models, Kimi K3 has the highest position at global rank 1. Its Category Score is 68.9, with an 80% uncertainty range of ±7.3. 55 visible models have a formal rank. This result applies only to this score run.
Available leaderboard data
- Models shown
- 100
- Formally ranked models
- 55
- Benchmark columns
- 26
- Dimensions with evidence
- 4/4
How to read the benchmark bars
Each bar compares models only within the same benchmark column. Bar lengths are relative to the models shown here; they are not Category Scores and cannot be compared across benchmark columns.
| Rank | Model | LMSpeed score | Planning & decomposition | Tool use | Environment & long-horizon execution | Recovery & completion reliability | Status | Evidence | Updated | ||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gert Labs50 models | DeepPlanning1 models | τ²-bench results84 models | MCP Atlas28 models | Toolathlon22 models | AA Tau3 Banking17 models | τ³-bench results9 models | MCP-Tasks5 models | GDPval-AA63 models | Terminal-Bench 2.051 models | BrowseComp29 models | OSWorld-Verified25 models | AA EnterpriseOps-Gym17 models | OSWorld 2.014 models | AA ITBench11 models | DeepSearchQA11 models | VITA-Bench9 models | WideResearch9 models | Claw-Eval28 models | APEX-Agents-AA24 models | JobBench22 models | ResearchClawBench19 models | AA Briefcase17 models | AA Harvey LAB16 models | CyberGym13 models | ExploitGym6 models | ||||||
| Formally ranked models55 | |||||||||||||||||||||||||||||||
| 1 | Kimi K3MoonshotAI | 68.9±7.3 | — | — | — | 84.2 | — | 33.4 | — | — | 1687.0 | 88.3 | 91.2 | — | 45.3 | — | 47.7 | 95.0 | — | — | — | 41.3 | 52.9 | — | 1542.0 | 94.6 | — | — | Rated | 3/4 dimensions · 12 families | Aug 9, 2026 |
| 2 | GPT-5.6 SolOpenAI | 68.2±7.3 | — | — | 85.1 | — | 58.0 | 33.0 | — | — | 1735.0 | 91.9 | 92.2 | — | 42.9 | 62.6 | 56.2 | — | — | — | — | — | — | — | 1504.0 | 87.2 | 84.5 | 33.7 | Rated | 3/4 dimensions · 12 families | Aug 9, 2026 |
| 3 | Claude Opus 5Anthropic | 67.5±8.2 | — | — | — | 85.8 | — | 30.3 | — | — | 1862.0 | — | 90.8 | — | 47.5 | 70.6 | — | 95.0 | — | — | — | — | — | — | 1720.0 | — | — | — | Rated | 3/4 dimensions · 8 families | Aug 9, 2026 |
| 4 | Claude Fable 5Anthropic | 65.7±8.2 | — | — | 98.5 | — | — | 26.8 | — | — | 1747.0 | 84.3 | — | 85.0 | 51.1 | — | — | — | — | — | — | — | — | — | 1574.0 | 93.6 | — | — | Rated | 3/4 dimensions · 7 families | Aug 9, 2026 |
| 5 | GPT-5.6 TerraOpenAI | 63.9±7.4 | — | — | 86.3 | — | 53.1 | 31.8 | — | — | 1583.0 | 87.4 | 87.5 | — | — | 50.2 | 51.0 | — | — | — | — | 38.9 | — | — | — | 85.2 | 81.8 | 23.2 | Rated | 3/4 dimensions · 11 families | Aug 9, 2026 |
| 6 | Claude Opus 4.8Anthropic | 63.2±5.3 | 73.0 | — | 94.4 | 82.2 | 59.9 | 27.6 | — | — | 1593.0 | 74.6 | 84.3 | 83.4 | 44.0 | 20.6 | — | 93.1 | — | — | — | — | — | 21.1 | 1345.0 | 91.1 | — | — | Rated | 4/4 dimensions · 13 families | Aug 9, 2026 |
| 7 | GPT-5.5OpenAI | 62.7±5.1 | 72.9 | — | 98.0 | 75.3 | 55.6 | — | — | — | 1490.0 | 82.0 | 84.4 | 78.7 | 46.6 | 13.0 | 45.8 | — | — | — | — | 37.7 | 42.7 | 17.0 | — | — | 81.8 | 13.4 | Rated | 4/4 dimensions · 15 families | Aug 9, 2026 |
| 8 | Grok 4.5SpaceXAI | 61.9±8.4 | — | — | — | — | — | 32.6 | — | — | 1527.0 | 83.3 | — | — | 40.8 | — | — | — | — | — | — | — | — | — | 1315.0 | 92.4 | — | — | Rated | 3/4 dimensions · 6 families | Aug 9, 2026 |
| 9 | Gemini 3.5 FlashGoogle | 61.1±5.6 | 61.9 | — | 95.3 | 83.6 | 56.5 | — | — | — | 1345.0 | 76.2 | — | 78.4 | 50.1 | — | — | — | — | — | — | 47.1 | — | 18.0 | — | — | — | — | Rated | 4/4 dimensions · 10 families | Aug 9, 2026 |
| 10 | GLM-5.2Z.ai | 60.5±7.2 | — | — | 99.1 | 76.8 | 48.2 | 26.8 | — | — | 1510.0 | 81.0 | — | — | 42.7 | — | 42.7 | — | — | — | — | 33.7 | — | 20.7 | 1253.0 | 91.0 | — | — | Rated | 3/4 dimensions · 11 families | Aug 9, 2026 |
| 11 | Muse Spark 1.1Meta | 60.4±7.1 | — | — | — | 88.1 | 75.6 | 25.2 | — | — | 1375.0 | 80.0 | — | 80.8 | 47.2 | 14.2 | — | 84.9 | — | — | — | — | 54.7 | — | 868.0 | 93.1 | 59.0 | 0.8 | Rated | 3/4 dimensions · 13 families | Aug 9, 2026 |
| 12 | GPT-5.6 LunaOpenAI | 59.8±7.4 | — | — | — | — | 53.4 | 27.2 | — | — | 1582.0 | 84.7 | 83.3 | — | — | 45.6 | 40.3 | — | — | — | — | 35.8 | — | — | — | 87.9 | 77.9 | 12.4 | Rated | 3/4 dimensions · 11 families | Aug 9, 2026 |
| 13 | Claude Sonnet 5Anthropic | 59.8±8.2 | — | — | — | — | — | 28.2 | — | — | 1603.0 | 80.4 | 84.7 | 81.2 | 44.7 | — | — | — | — | — | — | — | — | — | 1385.0 | 90.1 | — | — | Rated | 3/4 dimensions · 8 families | Aug 9, 2026 |
| 14 | Claude Opus 4.7 MaxAnthropic | 58±7.7 | — | — | 88.6 | 77.3 | — | — | — | — | 1491.0 | 69.4 | 79.3 | 78.0 | — | 18.2 | 46.7 | — | — | — | — | — | 45.9 | — | — | — | 73.1 | — | Rated | 3/4 dimensions · 9 families | Aug 9, 2026 |
| 15 | GPT-5.4OpenAI | 57.7±5.1 | 64.9 | — | 98.9 | 70.6 | 54.6 | — | — | — | 1391.0 | 75.1 | 82.7 | 75.0 | — | — | — | 73.6 | — | — | 60.3 | 33.3 | 38.9 | 15.3 | — | — | 79.0 | 6.0 | Rated | 4/4 dimensions · 15 families | Aug 9, 2026 |
| 16 | Qwen3.7 MaxQwen | 56.9±5.4 | 64.3 | — | 94.7 | 76.4 | — | — | — | — | 1270.0 | 69.7 | — | — | 45.0 | — | 42.5 | — | 47.9 | — | 65.2 | — | — | 18.7 | 914.0 | 83.4 | — | — | Rated | 4/4 dimensions · 12 families | Aug 9, 2026 |
| 17 | MiniMax M3MiniMax | 56.2±7.3 | — | — | 88.9 | 74.2 | — | — | — | — | 1390.0 | 66.0 | 83.5 | 70.1 | 32.1 | 4.6 | — | — | — | — | 74.5 | — | — | 19.8 | 1108.0 | 88.4 | — | — | Rated | 3/4 dimensions · 11 families | Aug 9, 2026 |
| 18 | Qwen3.6 27BQwen | 56.1±6.6 | 54.8 | — | 94.2 | — | — | — | — | — | 1138.0 | 59.3 | — | — | — | — | — | — | — | — | 72.4 | — | — | — | — | — | — | — | Rated | 4/4 dimensions · 5 families | Aug 9, 2026 |
| 19 | Kimi K2.6MoonshotAI | 55.5±5.3 | 56.8 | — | 95.9 | 55.9 | 50.0 | — | — | — | 1188.0 | 66.7 | 83.2 | 73.1 | — | 4.6 | — | 92.5 | — | 80.8 | 62.3 | 28.5 | — | 18.0 | — | — | — | — | Rated | 4/4 dimensions · 13 families | Aug 9, 2026 |
| 20 | Claude Opus 4.6Anthropic | 55.4±5.7 | 61.9 | — | 84.8 | — | — | — | — | — | — | 65.4 | 83.7 | 72.7 | — | — | — | 73.7 | — | — | 70.4 | 33.0 | 36.7 | 19.9 | — | — | 66.6 | — | Rated | 4/4 dimensions · 11 families | Aug 9, 2026 |
| 21 | Claude Opus 4.7Anthropic | 54.9±6.9 | 65.6 | — | 74.0 | — | — | — | — | — | — | — | — | — | — | 13.9 | — | — | — | — | — | — | — | 20.7 | — | — | — | — | Rated | 4/4 dimensions · 4 families | Aug 9, 2026 |
| 22 | Step 3.7 FlashStepFun | 54.1±5.7 | 51.6 | — | 98.5 | — | 49.5 | — | — | — | 1017.0 | 59.5 | 75.8 | — | — | — | — | 92.8 | — | — | 67.1 | 14.8 | — | — | — | — | — | — | Rated | 4/4 dimensions · 9 families | Aug 9, 2026 |
| 23 | Qwen3.7 PlusQwen | 54±5.7 | — | 62.3 | 93.0 | 73.2 | — | — | — | — | 943.0 | 70.3 | — | 73.3 | — | 2.8 | — | — | 45.6 | — | 62.7 | 22.4 | — | — | — | — | — | — | Rated | 4/4 dimensions · 9 families | Aug 9, 2026 |
| 24 | GLM-5.1Z.ai | 53.9±5.7 | 60.1 | — | 97.7 | 71.8 | — | — | 70.6 | — | 1256.0 | 63.5 | 68.0 | — | — | — | — | — | — | — | 62.3 | — | — | 18.2 | — | — | 68.7 | — | Rated | 4/4 dimensions · 9 families | Aug 9, 2026 |
| 25 | Claude Sonnet 4.6Anthropic | 53.2±6.1 | 62.9 | — | 79.5 | — | — | — | — | — | — | 59.1 | — | 72.1 | — | 8.3 | — | — | — | — | 67.8 | — | 36.9 | — | — | — | 65.2 | — | Rated | 4/4 dimensions · 7 families | Aug 9, 2026 |
| 26 | Gemini 3.6 FlashGoogle | 53±9.0 | — | — | — | — | — | 24.5 | — | — | 1423.0 | — | — | 83.0 | — | — | — | — | — | — | — | — | — | — | 962.0 | — | — | — | Rated | 3/4 dimensions · 4 families | Aug 9, 2026 |
| 27 | GPT-5.3 CodexOpenAI | 53±6.6 | 57.5 | — | 86.0 | — | — | — | — | — | — | 77.3 | — | 64.7 | — | — | — | — | — | — | — | — | 33.7 | — | — | — | — | — | Rated | 4/4 dimensions · 5 families | Aug 9, 2026 |
| 28 | GPT-5.4 MiniOpenAI | 51.8±8.1 | — | — | 93.4 | 57.7 | 42.9 | — | — | — | 1170.0 | 60.0 | — | 72.1 | — | — | — | — | — | — | — | 28.2 | — | — | — | — | — | — | Rated | 3/4 dimensions · 7 families | Aug 9, 2026 |
| 29 | DeepSeek V4 ProDeepSeek | 51.4±5.1 | 50.3 | — | 94.2 | 69.4 | 46.3 | 25.8 | — | — | 1293.0 | 59.1 | 80.4 | — | 40.4 | — | 38.3 | — | — | — | 59.8 | 24.3 | — | 17.1 | 931.0 | 84.4 | — | — | Rated | 4/4 dimensions · 14 families | Aug 9, 2026 |
| 30 | MiMo-V2.5-ProXiaomi | 51.1±5.9 | 62.7 | — | 94.2 | — | — | — | 72.9 | — | 1265.0 | 68.4 | — | — | — | — | 38.2 | — | — | — | 63.8 | 2.4 | — | — | 879.0 | 73.3 | — | — | Rated | 4/4 dimensions · 9 families | Aug 9, 2026 |
| 31 | MiMo-V2.5Xiaomi | 50.8±8.9 | 46.9 | — | — | — | — | — | — | — | — | 65.8 | — | — | — | — | — | — | — | — | 62.3 | — | — | 16.9 | — | — | — | — | Rated | 3/4 dimensions · 4 families | Aug 9, 2026 |
| 32 | Gemini 3.1 ProGoogle | 50.6±6.2 | 56.9 | — | 95.6 | — | — | — | — | — | 965.0 | — | — | — | — | — | — | 69.7 | — | — | 57.8 | 32.0 | — | 13.3 | — | — | — | — | Rated | 4/4 dimensions · 7 families | Aug 9, 2026 |
| 33 | DeepSeek V4 FlashDeepSeek | 50.3±5.7 | 54.4 | — | — | 64.0 | 40.7 | 31.1 | — | — | 1189.0 | 49.1 | 53.5 | — | — | — | — | — | — | — | 57.8 | — | — | — | — | — | 76.7 | — | Rated | 4/4 dimensions · 9 families | Aug 9, 2026 |
| 34 | Qwen3.6 PlusQwen | 50±5.5 | 50.6 | — | 97.7 | 48.2 | 39.8 | — | 70.7 | 74.1 | 1138.0 | 61.6 | — | — | — | — | — | — | 44.3 | 74.3 | 58.8 | — | — | 18.0 | — | — | — | — | Rated | 4/4 dimensions · 11 families | Aug 9, 2026 |
| 35 | InklingThinking Machines | 49.4±8.2 | — | — | — | 74.1 | — | 23.7 | — | — | 1238.0 | 63.8 | 77.1 | — | 38.1 | — | — | — | — | — | — | — | — | — | 840.0 | — | — | — | Rated | 3/4 dimensions · 7 families | Aug 9, 2026 |
| 36 | Claude Opus 4.5Anthropic | 49.3±5.3 | 64.2 | — | 86.3 | 42.3 | 43.5 | — | 70.2 | 71.8 | — | 59.3 | — | 66.3 | — | — | — | — | 23.3 | 76.4 | 59.6 | — | 32.3 | — | — | — | 50.6 | — | Rated | 4/4 dimensions · 12 families | Aug 9, 2026 |
| 37 | Grok 4.3SpaceXAI | 49.1±6.6 | 43.9 | — | 97.7 | — | — | — | — | — | 1084.0 | — | — | — | — | — | — | — | — | — | — | 17.0 | — | 12.4 | — | — | — | — | Rated | 4/4 dimensions · 5 families | Aug 9, 2026 |
| 38 | MiMo-V2-Pro | 48±8.9 | 36.7 | — | 95.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 57.8 | — | — | 15.3 | — | — | — | — | Rated | 3/4 dimensions · 4 families | Aug 9, 2026 |
| 39 | GPT-5.2OpenAI | 47.9±6.6 | 46.5 | — | 84.8 | — | — | — | — | — | — | — | 65.8 | 47.3 | — | — | — | — | — | — | — | — | 34.3 | — | — | — | — | — | Rated | 4/4 dimensions · 5 families | Aug 9, 2026 |
| 40 | GLM-4.7Z.ai | 47.1±8.6 | 40.0 | — | 95.9 | — | — | — | — | — | 1165.0 | 41.0 | 52.0 | — | — | — | — | — | 15.5 | — | — | — | — | — | — | — | — | — | Rated | 3/4 dimensions · 6 families | Aug 9, 2026 |
| 41 | Claude Sonnet 4.5Anthropic | 46.9±8.7 | 48.5 | — | — | — | — | — | — | — | — | 50.0 | — | 61.4 | — | — | — | — | 17.0 | — | — | — | 27.7 | — | — | — | — | — | Rated | 3/4 dimensions · 5 families | Aug 9, 2026 |
| 42 | Qwen3.6 35B A3BQwen | 46.8±5.9 | 42.6 | — | 95.3 | 62.8 | 26.9 | — | 67.2 | — | 1053.0 | 51.5 | — | — | — | — | — | — | 35.6 | 60.1 | 68.7 | — | — | — | — | — | — | — | Rated | 4/4 dimensions · 9 families | Aug 9, 2026 |
| 43 | Qwen3.5-27BQwen | 46.2±8.7 | 39.4 | — | 93.9 | — | — | — | — | — | — | 41.6 | 61.0 | 56.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Rated | 3/4 dimensions · 5 families | Aug 9, 2026 |
| 44 | Mistral Medium 3.5Mistral | 45.9±6.3 | 39.1 | — | 94.2 | — | — | — | 91.4 | — | 932.0 | — | — | — | 33.7 | — | — | — | — | — | — | — | — | — | 516.0 | 69.1 | — | — | Rated | 4/4 dimensions · 6 families | Aug 9, 2026 |
| 45 | GPT-5.4 NanoOpenAI | 45.5±8.1 | — | — | 92.5 | 56.1 | 35.5 | — | — | — | 1101.0 | 46.3 | — | 39.0 | — | — | — | — | — | — | — | 24.9 | — | — | — | — | — | — | Rated | 3/4 dimensions · 7 families | Aug 9, 2026 |
| 46 | Qwen3.5 | 45.4±5.3 | 46.8 | — | 95.6 | 46.1 | 36.3 | — | 68.4 | 74.2 | 963.0 | 52.5 | 62.0 | — | — | — | — | — | 43.7 | 74.0 | 56.8 | 15.3 | — | 14.2 | — | — | — | — | Rated | 4/4 dimensions · 13 families | Aug 9, 2026 |
| 47 | MiniMax M2.7MiniMax | 45.2±6.0 | 40.4 | — | 84.8 | — | 46.3 | — | — | — | 1158.0 | 57.0 | — | — | — | — | — | — | — | — | 48.7 | 10.6 | — | — | — | — | — | — | Rated | 4/4 dimensions · 7 families | Aug 9, 2026 |
| 48 | Gemini 3 FlashGoogle | 43.1±8.9 | 56.6 | — | 43.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 49.2 | — | 11.4 | — | — | — | — | — | Rated | 3/4 dimensions · 4 families | Aug 9, 2026 |
| 49 | Qwen3.5-35B-A3BQwen | 43.1±8.7 | 29.0 | — | 89.2 | — | — | — | — | — | — | 40.5 | 61.0 | 54.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Rated | 3/4 dimensions · 5 families | Aug 9, 2026 |
| 50 | Gemini 3.1 Flash LiteGoogle | 42.3±6.9 | 38.5 | — | 31.3 | — | — | — | — | — | 647.0 | — | — | — | — | — | — | — | — | — | — | 12.2 | — | — | — | — | — | — | Rated | 4/4 dimensions · 4 families | Aug 9, 2026 |
| 51 | GLM-5Z.ai | 41.7±5.6 | 51.0 | — | 98.2 | 31.1 | 38.0 | — | 65.6 | 60.8 | — | 56.2 | — | — | — | — | — | — | — | 69.8 | 57.7 | 14.5 | — | — | — | — | 43.2 | — | Rated | 4/4 dimensions · 10 families | Aug 9, 2026 |
| 52 | Gemini 3.5 Flash-LiteGoogle | 41.3±8.8 | — | — | — | — | — | 16.5 | — | — | 1139.0 | 54.0 | — | 74.0 | — | — | — | — | — | — | — | — | — | — | 635.0 | — | — | — | Rated | 3/4 dimensions · 5 families | Aug 9, 2026 |
| 53 | Kimi K2.5MoonshotAI | 39.7±5.1 | 45.9 | — | 95.9 | 29.5 | 27.8 | — | 65.7 | 59.1 | 1003.0 | 50.8 | 60.6 | — | — | — | — | 77.1 | — | 72.7 | 52.3 | 11.5 | 8.7 | 14.0 | — | — | — | — | Rated | 4/4 dimensions · 14 families | Aug 9, 2026 |
| 54 | DeepSeek V3.2DeepSeek | 39.2±6.9 | 29.6 | — | 78.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | 18.5 | — | 40.2 | — | — | — | — | — | — | — | Rated | 4/4 dimensions · 4 families | Aug 9, 2026 |
| 55 | gpt-oss-120bOpenAI | 35.3±6.2 | 29.6 | — | 65.8 | — | — | — | — | — | 802.0 | — | — | — | 25.5 | — | 5.6 | — | — | — | — | 3.1 | — | — | — | 13.9 | — | — | Rated | 4/4 dimensions · 7 families | Aug 9, 2026 |
| Estimated models — unranked31 | |||||||||||||||||||||||||||||||
| — | Qwen3.8 MaxQwen | 65.8±11.3 | — | — | — | — | — | — | — | — | — | — | — | 86.1 | — | 19.4 | — | — | — | 81.9 | — | — | 53.4 | — | — | — | — | — | Estimated | 2/4 dimensions · 3 families | Aug 9, 2026 |
| — | Inkling SmallThinking Machines | 55.8±10.9 | — | — | — | 79.6 | — | — | — | — | 1268.0 | 64.7 | 77.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 4 families | Aug 9, 2026 |
| — | Kimi K2.7 CodeMoonshotAI | 54±11.2 | — | — | 90.1 | 76.0 | — | — | — | — | 1189.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 3 families | Aug 9, 2026 |
| — | Qwen3.6 Max PreviewQwen | 54±11.8 | — | — | 95.9 | — | — | — | — | — | — | 65.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | GLM-5 TurboZ.ai | 52±11.9 | — | — | 98.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 55.8 | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | GPT-5.2 CodexOpenAI | 51.4±9.3 | 51.8 | — | 92.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 26.0 | — | — | — | — | — | Estimated | 3/4 dimensions · 3 families | Aug 9, 2026 |
| — | Ling-3.0-flashInclusionai | 50.5±10.2 | — | — | — | 65.5 | — | 28.0 | — | — | 1107.0 | — | 72.2 | — | — | — | — | — | — | 73.6 | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 5 families | Aug 9, 2026 |
| — | GPT-5.1 CodexOpenAI | 49.9±9.3 | 49.7 | — | 83.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 26.2 | — | — | — | — | — | Estimated | 3/4 dimensions · 3 families | Aug 9, 2026 |
| — | Gemini 3 ProGoogle | 49.4±9.3 | 63.2 | — | 87.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 11.4 | — | — | — | — | — | Estimated | 3/4 dimensions · 3 families | Aug 9, 2026 |
| — | Qwen3.5-122B-A10BQwen | 48.5±10.6 | — | — | 93.6 | — | — | — | — | — | 982.0 | 49.4 | 63.8 | 58.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 5 families | Aug 9, 2026 |
| — | GPT-5.1OpenAI | 47.8±9.3 | 41.2 | — | 81.9 | — | — | — | — | — | 988.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 3/4 dimensions · 3 families | Aug 9, 2026 |
| — | Qwen3 MaxQwen | 47.6±11.8 | 43.7 | — | 74.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | MiMo-V2-Flash | 47.3±11.8 | — | — | 83.9 | — | — | — | — | — | 838.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | GLM-5V TurboZ.ai | 47.1±9.3 | 30.8 | — | 98.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 53.8 | — | — | — | — | — | — | — | Estimated | 3/4 dimensions · 3 families | Aug 9, 2026 |
| — | Hy3 previewTencent | 46.5±11.2 | 36.9 | — | — | — | — | — | — | — | 1215.0 | 54.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 3 families | Aug 9, 2026 |
| — | Command ACohere | 46.1±11.8 | — | — | 85.0 | — | — | — | — | — | 718.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | GPT-5OpenAI | 45.6±9.3 | — | — | 86.5 | — | — | — | — | — | 1082.0 | — | — | — | — | — | — | — | — | — | — | — | 8.5 | — | — | — | — | — | Estimated | 3/4 dimensions · 3 families | Aug 9, 2026 |
| — | Claude Sonnet 4Anthropic | 44.5±9.3 | 39.7 | — | 52.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 18.4 | — | — | — | — | — | Estimated | 3/4 dimensions · 3 families | Aug 9, 2026 |
| — | Ling 2.6 FlashinclusionAI | 44.4±11.8 | — | — | 86.0 | — | — | — | — | — | 550.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | Trinity Large ThinkingArcee AI | 43.9±9.3 | 32.5 | — | 90.1 | — | — | — | — | — | 564.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 3/4 dimensions · 3 families | Aug 9, 2026 |
| — | Gemini 2.5 ProGoogle | 43.7±9.3 | 42.0 | — | 54.1 | — | — | — | — | — | 669.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 3/4 dimensions · 3 families | Aug 9, 2026 |
| — | Grok 4.20SpaceXAI | 43.3±11.2 | 38.4 | — | — | — | — | — | — | — | — | 47.1 | — | — | — | — | — | 62.8 | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 3 families | Aug 9, 2026 |
| — | MiMo-V2-Omni | 41.5±11.9 | — | — | 91.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 45.2 | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | GPT-4.1 MiniOpenAI | 40.4±11.8 | — | — | 52.9 | — | — | — | — | — | 505.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | GPT-4.1OpenAI | 39.9±11.8 | 25.6 | — | 47.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | Mistral Large 3 | 39.4±11.8 | — | — | 24.6 | — | — | — | — | — | 640.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | gpt-oss-20bOpenAI | 37.7±9.3 | — | — | 60.2 | — | — | — | — | — | 564.0 | — | — | — | — | — | — | — | — | — | — | 0.7 | — | — | — | — | — | — | Estimated | 3/4 dimensions · 3 families | Aug 9, 2026 |
| — | DeepSeek V3 | 34.6±11.8 | — | — | 22.8 | — | — | — | — | — | 231.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | Llama 4 MaverickMeta | 33.3±11.8 | — | — | 17.8 | — | — | — | — | — | 5.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | GPT-4.1 NanoOpenAI | 33.2±11.8 | — | — | 17.3 | — | — | — | — | — | 62.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| — | Llama 4 ScoutMeta | 32.9±11.8 | — | — | 15.5 | — | — | — | — | — | 111.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Estimated | 2/4 dimensions · 2 families | Aug 9, 2026 |
| Provisional models — unranked14 | |||||||||||||||||||||||||||||||
| — | Muse Spark 1.2Meta | 62.3±16.0 | — | — | — | — | — | — | — | — | 1631.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | GPT-5.5 ProOpenAI | 61.4±16.1 | — | — | — | — | — | — | — | — | — | — | 90.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | GPT-5.4 ProOpenAI | 60.5±16.1 | — | — | — | — | — | — | — | — | — | — | 89.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | Laguna S 2.1Poolside | 53.7±16.0 | — | — | — | — | — | — | — | — | — | 70.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | Hy3Tencent | 52.8±16.0 | — | — | — | — | — | — | — | — | 1215.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | GPT-5.1 Codex MaxOpenAI | 50.1±16.0 | — | — | 83.0 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | Grok Build 0 1SpaceXAI | 50±16.1 | 49.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | O3OpenAI | 49.5±16.0 | — | — | 80.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | GLM-4.6Z.ai | 48.6±16.0 | — | — | 76.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | Claude Opus 4.1Anthropic | 46.8±16.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 21.9 | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | O1OpenAI | 45.8±16.0 | — | — | 62.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | Kimi K2MoonshotAI | 45.5±16.0 | — | — | 61.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | Qwen3.5 PlusAlibaba | 44.7±16.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 18.5 | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
| — | GLM-4.5 AirZ.ai | 43.1±16.0 | — | — | 46.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | Provisional | 1/4 dimensions · 1 families | Aug 9, 2026 |
What this leaderboard measures
Which AI model is better suited to agent tasks?
This leaderboard looks at whether a model can plan, use tools, act in an environment, and recover from failure. It measures more than the ability to answer a question.
Four capability dimensions
The four dimensions come from the category blueprint. Available data may cover only some of them. A dimension without evidence is not presented as a verified capability.
Planning & decomposition
2 benchmark columns currently provide evidence for this dimension.
- Gert Labs
- DeepPlanning
Tool use
6 benchmark columns currently provide evidence for this dimension.
- τ²-bench results
- MCP Atlas
- Toolathlon
- AA Tau3 Banking
- τ³-bench results
- MCP-Tasks
Environment & long-horizon execution
10 benchmark columns currently provide evidence for this dimension.
- GDPval-AA
- Terminal-Bench 2.0
- BrowseComp
- OSWorld-Verified
- AA EnterpriseOps-Gym
- OSWorld 2.0
- AA ITBench
- DeepSearchQA
- VITA-Bench
- WideResearch
Recovery & completion reliability
8 benchmark columns currently provide evidence for this dimension.
- Claw-Eval
- APEX-Agents-AA
- JobBench
- ResearchClawBench
- AA Briefcase
- AA Harvey LAB
- CyberGym
- ExploitGym
Agent tasks this page can help with
- Research assistants that search, organize evidence, and adjust a plan when new facts appear.
- Browser and computer tasks that require several tool calls and interface actions.
- Multi-step business automation that must follow rules, handle failure, and leave reviewable output.
How to choose a model with this leaderboard
- Step 1
Check the rating status first
Only Rated models receive a rank. Estimated and Provisional models do not have a formal position.
- Step 2
Review uncertainty and evidence
When scores are close, do not rely on rank alone. Check uncertainty, dimension coverage, and benchmark count.
- Step 3
Test the real task last
A leaderboard cannot replace your own test. Check quality, speed, price, context, and provider limits together.
Start with your tools and execution environment. Then compare recovery, speed, cost, and permission controls. The top-ranked model is not best for every agent.
Rating status guide
Rated
Rated means the evidence and overlap rules are met. The model can receive a formal rank.
Estimated
Estimated means there is useful evidence, but it is not enough for a formal rank.
Provisional
Provisional means evidence is limited or dimension and benchmark-family coverage is below the estimated threshold. Use the result only as an early signal.
Benchmarks and evidence sources
Evidence source names and benchmark groups come from the currently available score data. One source may contribute several benchmarks.
BenchLM
AA Briefcase, AA EnterpriseOps-Gym, AA Harvey LAB, AA ITBench, AA Tau3 Banking, APEX-Agents-AA, BrowseComp, Claw-Eval, CyberGym, DeepPlanning, DeepSearchQA, ExploitGym, GDPval-AA, Gert Labs, JobBench, MCP Atlas, MCP-Tasks, OSWorld 2.0, OSWorld-Verified, ResearchClawBench, Terminal-Bench 2.0, Toolathlon, VITA-Bench, WideResearch, τ²-bench results, and τ³-bench results
How the AI agent model ranking is built
LMSpeed combines eligible third-party benchmarks inside four fixed capability dimensions. Rated models meet the evidence and overlap requirements for a formal rank; Estimated and Provisional models remain visible without receiving a rank.
Read the Category Score methodologyLeaderboard limits
Category Scores use the third-party benchmarks currently included by LMSpeed. Tests can use different data, prompts, and scoring rules. The result is not permanent and cannot represent every real task. Test important choices with your own data and workflow.
Frequently asked questions
Which visible model has the highest formal rank now?
Among the currently visible models, Kimi K3 has the highest formal position at global rank 1. Its Category Score is 68.9. 55 visible models meet the formal ranking rules. This result applies only to the run date and methodology version shown on the page.
Can I compare scores across different categories?
No. Each category uses different capability dimensions and evidence. A Category Score is comparable only inside the same leaderboard. Review the matching category for each task.
Are Estimated and Provisional models still useful?
They can help you find candidates, but their evidence is not complete enough for a formal rank. Review coverage and uncertainty, then test the model on a real task.
How often does the leaderboard update?
The leaderboard updates after a new completed score run is published. The current run date and methodology version appear above. LMSpeed does not promise a fixed daily or weekly schedule.
Is the number one model always best for me?
No. Your result also depends on speed, price, context length, tool support, region, and provider limits. Use the leaderboard to narrow the field, then run your own test.
How is the agent ranking different from the reasoning ranking?
The reasoning ranking focuses on logic, causality, and evidence checks. The agent ranking also covers planning, tool use, environment execution, and recovery. Strong reasoning alone does not prove that a model can finish an agent workflow.
