Choose a model to open its comparison page.
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of....
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
Input and output token limits for this model, plus how it ranks on long-context understanding.
1 metric
2 metrics
2 metrics
5 metrics · Estimated
6 metrics · Rated
5 metrics · Estimated
3 metrics · Provisional
3 metrics · Provisional
0 metrics · No data
4 metrics · Estimated
2 metrics · Provisional
Third-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
DeepInfra deepinfra/fp8 | $0.500/M | $1.20/M | 99.9% | — | — | 524.3K tokens / 262.1K tokens |
Together together | $0.500/M | $1.20/M | 99.7% | — | — | 524.3K tokens / — |
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
gpt-5-3-codex
OpenAI GPT-5.3 Codex is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks.
gpt-5-2
OpenAI GPT-5.2 is a GPT-5 series model emphasizing advanced reasoning, multimodal understanding, and high-quality outputs for complex enterprise workloads.
gpt-5-4-mini
OpenAI GPT-5.4 Mini is a compact language model in the GPT-5 series, optimized for quick responses and high throughput.
claude-opus-4-6
Anthropic Claude Opus 4.6 is the most capable Claude Opus tier, optimized for complex analysis, long-horizon coding, and high-stakes enterprise reasoning workloads.
claude-sonnet-4-6
Anthropic Claude Sonnet 4.6 extends the Sonnet line with improved tool use, coding reliability, and long-context performance for everyday production workloads.