Choose a model to open its comparison page.
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and ...
Оценки наблюдаемой способности; неопределённость независимых семейств бенчмарков показана отдельно.
Input and output token limits for this model, plus how it ranks on long-context understanding.
1 метрика
2 метрики
2 метрики
4 метрики · Оценка
5 метрик · Предварительная
3 метрики · Оценка
6 метрик · Предварительная
0 метрик · Нет данных
0 метрик · Нет данных
0 метрик · Нет данных
3 метрики · Предварительная
Сторонние данные эндпоинтов OpenRouter показаны отдельно от измерений LMSpeed. Некоторые 30-минутные live-метрики появляются только после синхронизации с OpenRouter API key.
| Provider endpoint | Вход | Вывод | Доступность 1 день | Задержка 30м | Пропускная способность 30м | Контекст / вывод |
|---|---|---|---|---|---|---|
Novita novita | $0.010/M | $0.030/M | 99.3% | — | — | 262.1K токенов / 32.8K токенов |
glm-4-7-flash
Zhipu AI GLM-4.7 Flash is a speed-focused GLM variant for real-time chat, function calling, and bilingual enterprise copilots with competitive token economics.
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
deepseek-v3-2
DeepSeek V3.2 is an upgraded V3-series MoE model with stronger reasoning, coding, and math performance, widely available through OpenAI-compatible API relays.
glm-5
Zhipu GLM-5 is Zhipu flagship GLM series model with enhanced reasoning, agent capabilities, and strong performance on Chinese enterprise and coding scenarios.
minimax-m2-7
MiniMax M2.7 is a high-tier M2-series model tuned for complex reasoning, long-context dialogue, and production-grade API workloads.