Category Score V3 leaderboard

LMSpeed Best Multilingual Models

Compare multilingual AI models across cross-language understanding, generation, reasoning transfer, and low-resource robustness benchmarks using a dimension-balanced score.

Methodology 3.0Methodology

Current answer

No model currently meets the formal ranking rules. The table has 7 Estimated models and 10 Provisional models. They are useful signals, but they are not formal ranks.

Rankings can change when data or methods change. The run date appears above.

Available leaderboard data

Models shown
17
Formally ranked models
0
Benchmark columns
4
Dimensions with evidence
2/4

How to read the benchmark bars

Each bar compares models only within the same benchmark column. Bar lengths are relative to the models shown here; they are not Category Scores and cannot be compared across benchmark columns.

RankModelLMSpeed scoreCross-language understandingReasoning transferStatusEvidenceUpdated
NOVA-63undefined modelsAA Global-MMLU-Liteundefined modelsINCLUDEundefined modelsMMLU-ProXundefined models
Estimated models — unranked7
Qwen3.7 MaxQwen
58.1±11.7
59.086.287.0Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.5
54.6±12.2
59.184.7Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.6 PlusQwen
52.2±12.2
57.984.7Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4.5Anthropic
51.7±12.2
56.785.7Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.7 PlusQwen
51.1±11.7
58.883.085.4Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Kimi K2.5MoonshotAI
44.2±12.2
56.082.3Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
GLM-5Z.ai
43.8±12.2
55.183.1Estimatedundefined/4 dimensions · undefined familiesAug 28, 2026
Provisional models — unranked10
Claude Opus 5Anthropic
64.3±17.1
89.8Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 3.1 ProGoogle
59.7±17.3
93.2Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4.8Anthropic
55±17.1
87.6Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Claude Opus 4.6Anthropic
54±17.3
92.2Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Gemini 3 ProGoogle
54±17.3
92.2Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.5-27BQwen
45.4±16.4
82.2Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.5-122B-A10BQwen
45.4±16.4
82.2Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3.5-35B-A3BQwen
41.6±16.4
81.0Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
Qwen3
36.7±16.4
79.4Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026
gpt-oss-120bOpenAI
31.5±17.3
82.8Provisionalundefined/4 dimensions · undefined familiesAug 28, 2026

What this leaderboard measures

Which AI model is better suited to multilingual work?

This leaderboard covers cross-language understanding, multilingual generation, reasoning transfer, and low-resource language performance. A current run may cover only some of these dimensions.

Available data covers 2/4 dimensions and shows 4 benchmark columns.

Four capability dimensions

The four dimensions come from the category blueprint. Available data may cover only some of them. A dimension without evidence is not presented as a verified capability.

Cross-language understanding

3 benchmark columns currently provide evidence for this dimension.

  • NOVA-63
  • AA Global-MMLU-Lite
  • INCLUDE

Multilingual generation

Available data has no eligible evidence for this dimension, so it does not affect this Category Score.

Reasoning transfer

1 benchmark columns currently provide evidence for this dimension.

  • MMLU-ProX

Low-resource robustness

Available data has no eligible evidence for this dimension, so it does not affect this Category Score.

Multilingual tasks this page can help with

  • Translation and localization that preserve meaning, tone, and specialist terms.
  • Multilingual support that understands varied phrasing and stays in the target language.
  • Cross-language research that finds material in one language and answers in another.

How to choose a model with this leaderboard

  1. Step 1

    Check the rating status first

    Only Rated models receive a rank. Estimated and Provisional models do not have a formal position.

  2. Step 2

    Review uncertainty and evidence

    When scores are close, do not rely on rank alone. Check uncertainty, dimension coverage, and benchmark count.

  3. Step 3

    Test the real task last

    A leaderboard cannot replace your own test. Check quality, speed, price, context, and provider limits together.

Test the exact target language. Check dialects, writing systems, cultural phrasing, and specialist vocabulary.

Rating status guide

Rated

Rated means the evidence and overlap rules are met. The model can receive a formal rank.

Estimated

Estimated means there is useful evidence, but it is not enough for a formal rank.

Provisional

Provisional means evidence is limited or dimension and benchmark-family coverage is below the estimated threshold. Use the result only as an early signal.

Benchmarks and evidence sources

Evidence source names and benchmark groups come from the currently available score data. One source may contribute several benchmarks.

  • BenchLM

    AA Global-MMLU-Lite, INCLUDE, MMLU-ProX, and NOVA-63

How the multilingual model ranking is built

LMSpeed combines eligible third-party benchmarks inside four fixed capability dimensions. Rated models meet the evidence and overlap requirements for a formal rank; Estimated and Provisional models remain visible without receiving a rank.

Read the Category Score methodology

Leaderboard limits

Category Scores use the third-party benchmarks currently included by LMSpeed. Tests can use different data, prompts, and scoring rules. The result is not permanent and cannot represent every real task. Test important choices with your own data and workflow.

Frequently asked questions

Which visible model has the highest formal rank now?

There is no formal number one now. The page has 7 Estimated models and 10 Provisional models. They do not have a formal rank and should not be called the winner.

Can I compare scores across different categories?

No. Each category uses different capability dimensions and evidence. A Category Score is comparable only inside the same leaderboard. Review the matching category for each task.

Are Estimated and Provisional models still useful?

They can help you find candidates, but their evidence is not complete enough for a formal rank. Review coverage and uncertainty, then test the model on a real task.

How often does the leaderboard update?

The leaderboard updates after a new completed score run is published. The current run date and methodology version appear above. LMSpeed does not promise a fixed daily or weekly schedule.

Is the number one model always best for me?

No. Your result also depends on speed, price, context length, tool support, region, and provider limits. Use the leaderboard to narrow the field, then run your own test.

Does a high multilingual score mean every language is supported?

No. The score reflects only the language evidence in the current run. Test each language, dialect, writing system, and specialist domain separately. Low-resource languages need real examples.