Inkling Small API Benchmarks, Pricing & Provider Data
Compare Inkling Small with another model
Choose a model to open its comparison page.
Inkling Small benchmark, API pricing, and provider data cover undefined API provider} other undefined API providers}}, with prices starting at $0.00005/request.
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of....
Category Performance
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
- Coverage
- 7 / 8
- Methodology
- V3.0
#1Instruction following57.6Provisional1/4 Measured dimensions
#2Agents54.6Estimated2/4 Measured dimensions
#3Math52.5Provisional1/4 Measured dimensions
#4Coding51.9RatedGlobal rank #214/4 Measured dimensions
#5Reasoning50.9RatedGlobal rank #383/4 Measured dimensions
#6Knowledge49.7Provisional1/4 Measured dimensions
#7Multimodal45.8Estimated2/4 Measured dimensions
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Released
- Jul 2026
- Tokenizer
- Other
- Architecture
- text+image+audio->text
- Moderated
- No
- Supported parameters
- frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyseedstoptemperaturetool_choicetoolstop_ktop_logprobstop_p
Rankings
Excels at
It's decent at
Falls behind in
Detailed scores
Updated: Sep 13, 2026Overall
undefined metric} other undefined metrics}}
Overall score56.0#65 / 112
Overall
undefined metric} other undefined metrics}}
Speed & latency
undefined metric} other undefined metrics}}
Output speed146.4 tok/s#27 / 80Time to first token1.85 s#51 / 80
Speed & latency
undefined metric} other undefined metrics}}
Pricing
undefined metric} other undefined metrics}}
Input price$0.300/M#55 / 186Output price$1.20/M#50 / 186
Pricing
undefined metric} other undefined metrics}}
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Score54.680% interval43.7–65.42/4 Measured dimensionsAgentic score56.7#48 / 77Terminal-Bench 2.064.7#26 / 53
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Coding
V3.0undefined metric} other undefined metrics}} · Rated
Score51.9#2180% interval45.0–58.74/4 Measured dimensionsSciCode49.7%#35 / 89Coding score52.6#47 / 87
Coding
V3.0undefined metric} other undefined metrics}} · Rated
Reasoning
V3.0undefined metric} other undefined metrics}} · Rated
Score50.9#3880% interval42.2–59.63/4 Measured dimensionsGPQA89.5%#41 / 218HLE33.3%#51 / 216
Reasoning
V3.0undefined metric} other undefined metrics}} · Rated
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Score49.780% interval35.8–63.71/4 Measured dimensionsKnowledge score66.0#47 / 83GPQA-D89.5#18 / 32
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Math
V3.0undefined metric} other undefined metrics}} · Provisional
Score52.580% interval38.1–66.91/4 Measured dimensionsMath score76.9#13 / 63AIME2695.5#6 / 14
Math
V3.0undefined metric} other undefined metrics}} · Provisional
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
Multimodal
V3.0undefined metric} other undefined metrics}} · Estimated
Score45.880% interval33.9–57.72/4 Measured dimensionsMultimodal Grounded score48.8#50 / 56MMMU-Pro74.0#26 / 31
Multimodal
V3.0undefined metric} other undefined metrics}} · Estimated
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
Score57.680% interval41.6–73.61/4 Measured dimensionsInstruction Following score89.6#19 / 52IFBench82.2#2 / 14
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
OpenRouter endpoints
3 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
Together together | $0.500/M | $1.20/M | 100% | — | — | undefined tokens / undefined tokens |
BaseTen baseten/fp8 | $0.500/M | $1.20/M | 100.0% | — | — | undefined tokens / undefined tokens |
DeepInfra deepinfra/fp8 | $0.450/M | $1.20/M | 99.4% | — | — | undefined tokens / undefined tokens |
Pricing Comparison
Compare Inkling Small API pricing across 12 providers. Prices range from $0.00005/request to $75.00/M. CM-API 公益站 offers the lowest rate at $0.00005/request.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% | thinkingmachines/inkling-small:free | default | $75.00/M | $75.00/M | — | — | — | |
L1 100% | thinkingmachines/inkling-small:free | free | $0.00005/request | - | — | — | — | |
L1 99% | thinkingmachines/inkling-small:free | default | $0.073/request | - | — | — | — | |
L1 100% | thinkingmachines/inkling-small:free | default | $0.010/request | - | — | — | — | |
L1 0% L2 100% | inkling-small | price | -100%$0.0003/M Cache read$0.00006/MCache write$0.0003/MCache write 1h$0.0005/M | -100%$0.0012/M | — | — | — | |
L1 0% | inkling-small | model | $1.46/M Cache read$0.0073/M | $14.60/M | — | — | — | |
L1 0% | inkling-small | 无限制 | $1.46/M Cache read$0.0073/M | $14.60/M | — | — | — |
Alternatives & Similar Models
DeepSeek V4 Flash
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
DeepSeek V4 Pro
deepseek-v4-pro
DeepSeek V4 Pro is the professional-tier DeepSeek V4 model, targeting frontier reasoning, coding, and agent workflows with maximum capability.
MiMo-V2.5-Pro
mimo-v2-5-pro
Xiaomi MiMo-V2.5-Pro is a large open-source language model in the MiMo series, offering advanced reasoning and general-purpose capabilities.
Kimi K3
kimi-k3
Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating...
MiMo-V2.5
mimo-v2-5
Xiaomi MiMo-V2.5 is a native omnimodal sparse MoE model (310B total, 15B active) with unified text, image, video, and audio understanding, built on the MiMo-V2-Flash backbone with dedicated vision and audio encoders. It supports up to 1M tokens of context, strong agentic workflows, and open weights on Hugging Face.
IInkling
inkling
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Frequently Asked Questions
- What benchmark data does Inkling Small include?
- LMSpeed shows Inkling Small benchmark context, API price, output speed, first-token latency, and provider data across 12 providers when those signals are available.
- What is the Inkling Small API price?
- Inkling Small has pricing from undefined provider} other undefined providers}}, ranging from $0.00005/request to $75.00/M. CM-API 公益站 has the lowest listed price.
- What does the Inkling Small API pricing table include?
- The Inkling Small API pricing table compares 12 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Inkling Small API pricing?
- CM-API 公益站 currently has the lowest listed Inkling Small price at $0.00005/request across undefined provider} other undefined providers}}.
- Is Inkling Small API free?
- Inkling Small does not currently have a free API tier on LMSpeed. All 12 providers charge per token.
