Category Score V3 leaderboard

LMSpeed Best Multimodal Models

Compare multimodal AI models across perception, OCR, document and spatial understanding, visual reasoning, video, and grounded-action benchmarks in one leaderboard.

Methodology 3.0Methodology

Current answer

Among the currently visible formally ranked models, Qwen3.8 Max has the highest position at global rank 1. Its Category Score is 64.7, with an 80% uncertainty range of ±6.3. 10 visible models have a formal rank. This result applies only to this score run.

View Qwen3.8 Max detailsRankings can change when data or methods change. The run date appears above.

Available leaderboard data

Models shown
47
Formally ranked models
10
Benchmark columns
10
Dimensions with evidence
4/4

How to read the benchmark bars

Each bar compares models only within the same benchmark column. Bar lengths are relative to the models shown here; they are not Category Scores and cannot be compared across benchmark columns.

RankModelLMSpeed scorePerception & OCRDocument & spatial understandingVisual reasoningVideo & grounded actionStatusEvidenceUpdated
V*undefined modelsSimpleVQAundefined modelsCharXivundefined modelsMMMU-Proundefined modelsMathVisionundefined modelsERQAundefined modelsMedXpertQA (MM)undefined modelsBenchLM Multimodal Grounded scoreundefined modelsScreenSpot Proundefined modelsVideoMMMUundefined models
Formally ranked models10
1Qwen3.8 MaxQwen
64.7±6.3
75.093.582.395.277.880.484.588.7Ratedundefined/4 dimensions · undefined familiesAug 28, 2026
2Gemini 3.1 ProGoogle
55.4±6.6
72.480.283.969.481.384.4Ratedundefined/4 dimensions · undefined familiesAug 28, 2026
3Qwen3.7 PlusQwen
55.1±6.3
81.785.979.090.369.871.079.085.4Ratedundefined/4 dimensions · undefined familiesAug 28, 2026
4GPT-5.4OpenAI
52.9±6.6
61.182.881.265.477.185.4Ratedundefined/4 dimensions · undefined familiesAug 28, 2026
5Qwen3.6 PlusQwen
50.7±6.5
96.981.578.888.068.284.0Ratedundefined/4 dimensions · undefined familiesAug 28, 2026
6Qwen3.5
50.2±6.5
95.880.879.088.665.684.7Ratedundefined/4 dimensions · undefined familiesAug 28, 2026
7Gemini 3 ProGoogle
49.9±6.5
88.081.481.086.672.787.6Ratedundefined/4 dimensions · undefined familiesAug 28, 2026
8Qwen3.6 27BQwen
46.3±6.5
94.756.178.475.862.584.4Ratedundefined/4 dimensions · undefined familiesAug 28, 2026
9Qwen3.6 35B A3BQwen
44.1±7.0
58.978.075.383.7Ratedundefined/4 dimensions · undefined familiesAug 28, 2026
10Claude Opus 4.5Anthropic
32.4±6.5
67.068.570.674.345.784.4Ratedundefined/4 dimensions · undefined familiesAug 28, 2026
Estimated models — unranked18
Kimi K3MoonshotAI
65±11.3
91.381.694.3Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4.8Anthropic
59.6±12.0
89.987.9Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.8 27BQwen
57.7±11.3
90.290.065.5Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 3.5 FlashGoogle
56.3±11.9
84.283.6Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Kimi K2.6MoonshotAI
53.5±9.0
96.980.479.487.4Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Kimi K2.5MoonshotAI
52.7±12.0
78.586.6Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
MiMo-V2.5Xiaomi
49.1±11.9
81.077.9Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.5-27BQwen
49±12.1
93.786.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4.6Anthropic
48.4±11.1
77.351.664.883.1Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.5-122B-A10BQwen
47.7±9.4
93.277.286.2Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
MiniMax M3MiniMax
46.9±12.0
78.184.6Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Muse Glimmer 30BMeta
46.4±9.4
78.874.075.4Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Inkling SmallThinking Machines
45.9±11.9
81.374.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.5-35B-A3BQwen
45.8±12.1
92.783.9Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
InklingThinking Machines
45.8±11.9
82.073.5Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.2OpenAI
42.3±9.0
75.982.179.583.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Grok 4.20SpaceXAI
40.2±8.9
57.460.975.254.165.8Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Command ACohere
29.1±11.9
52.763.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Provisional models — unranked19
GPT-5.4 ProOpenAI
73±16.1
94.0Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4.7 MaxAnthropic
60.7±16.1
91.0Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.6 SolOpenAI
59.5±16.1
83.0Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Step 3.7 FlashStepFun
59±14.5
95.379.2Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 3.7 FlashGoogle
57.2±16.1
88.7Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Muse Spark 1.1Meta
56.8±16.1
88.4Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Sonnet 5Anthropic
56.6±16.1
88.3Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 5Anthropic
56.4±17.3
85.9Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.5OpenAI
55.7±16.1
81.2Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.6 TerraOpenAI
54.7±16.1
80.7Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 3.6 FlashGoogle
51.1±17.3
72.7Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.6 LunaOpenAI
50.3±16.1
78.4Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Grok 4.3SpaceXAI
49.8±16.1
78.1Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.4 MiniOpenAI
47.1±16.1
76.6Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 3.5 Flash-LiteGoogle
46.9±17.3
62.3Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Fable 5Anthropic
46.3±17.3
60.8Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Sonnet 4.6Anthropic
45.7±16.1
77.4Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 3.1 Flash LiteGoogle
42.6±16.1
73.2Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
GPT-5.4 NanoOpenAI
31.2±16.1
66.1Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026

What this leaderboard measures

Which AI model is better suited to image and document tasks?

This leaderboard covers perception, OCR, document and spatial understanding, visual reasoning, video, and action. Supported input types can differ between models.

Available data covers 4/4 dimensions and shows 10 benchmark columns.

Four capability dimensions

The four dimensions come from the category blueprint. Available data may cover only some of them. A dimension without evidence is not presented as a verified capability.

Perception & OCR

2 benchmark columns currently provide evidence for this dimension.

  • V*
  • SimpleVQA

Document & spatial understanding

1 benchmark columns currently provide evidence for this dimension.

  • CharXiv

Visual reasoning

5 benchmark columns currently provide evidence for this dimension.

  • MMMU-Pro
  • MathVision
  • ERQA
  • MedXpertQA (MM)
  • BenchLM Multimodal Grounded score

Video & grounded action

2 benchmark columns currently provide evidence for this dimension.

  • ScreenSpot Pro
  • VideoMMMU

Multimodal tasks this page can help with

  • Document extraction that reads text, tables, layout, and page relationships.
  • Image questions that combine visual details with written reasoning.
  • Video or visual agents that understand a process and use visual information to act.

How to choose a model with this leaderboard

  1. Step 1

    Check the rating status first

    Only Rated models receive a rank. Estimated and Provisional models do not have a formal position.

  2. Step 2

    Review uncertainty and evidence

    When scores are close, do not rely on rank alone. Check uncertainty, dimension coverage, and benchmark count.

  3. Step 3

    Test the real task last

    A leaderboard cannot replace your own test. Check quality, speed, price, context, and provider limits together.

Confirm supported files and input types first. Also check OCR languages, file limits, speed, cost, and privacy.

Rating status guide

Rated

Rated means the evidence and overlap rules are met. The model can receive a formal rank.

Estimated

Estimated means there is useful evidence, but it is not enough for a formal rank.

Provisional

Provisional means evidence is limited or dimension and benchmark-family coverage is below the estimated threshold. Use the result only as an early signal.

Benchmarks and evidence sources

Evidence source names and benchmark groups come from the currently available score data. One source may contribute several benchmarks.

  • BenchLM

    BenchLM Multimodal Grounded score, CharXiv, ERQA, MathVision, MedXpertQA (MM), MMMU-Pro, ScreenSpot Pro, SimpleVQA, V*, and VideoMMMU

How the multimodal model ranking is built

LMSpeed combines eligible third-party benchmarks inside four fixed capability dimensions. Rated models meet the evidence and overlap requirements for a formal rank; Estimated and Provisional models remain visible without receiving a rank.

Read the Category Score methodology

Leaderboard limits

Category Scores use the third-party benchmarks currently included by LMSpeed. Tests can use different data, prompts, and scoring rules. The result is not permanent and cannot represent every real task. Test important choices with your own data and workflow.

Frequently asked questions

Which visible model has the highest formal rank now?

Among the currently visible models, Qwen3.8 Max has the highest formal position at global rank 1. Its Category Score is 64.7. 10 visible models meet the formal ranking rules. This result applies only to the run date and methodology version shown on the page.

Can I compare scores across different categories?

No. Each category uses different capability dimensions and evidence. A Category Score is comparable only inside the same leaderboard. Review the matching category for each task.

Are Estimated and Provisional models still useful?

They can help you find candidates, but their evidence is not complete enough for a formal rank. Review coverage and uncertainty, then test the model on a real task.

How often does the leaderboard update?

The leaderboard updates after a new completed score run is published. The current run date and methodology version appear above. LMSpeed does not promise a fixed daily or weekly schedule.

Is the number one model always best for me?

No. Your result also depends on speed, price, context length, tool support, region, and provider limits. Use the leaderboard to narrow the field, then run your own test.

Does a high multimodal score support every image and video format?

No. Benchmark scores and product feature support are different. Confirm file formats, size limits, video support, OCR languages, and provider APIs before choosing a model.