Inkling API Benchmarks, Pricing & Provider Data
Compare Inkling with another model
Choose a model to open its comparison page.
Inkling benchmark, API pricing, and provider data cover undefined API provider} other undefined API providers}}, with prices starting at $0.00000001/request. The page also shows measured API speed and first-token latency.
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic an...
Category Performance
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
- Coverage
- 7 / 8
- Methodology
- V3.0
#1Math65.4Provisional1/4 Measured dimensions
#2Instruction following56.4Provisional1/4 Measured dimensions
#3Reasoning55.2Estimated2/4 Measured dimensions
#4Knowledge51.3Provisional1/4 Measured dimensions
#5Agents48.6RatedGlobal rank #423/4 Measured dimensions
#6Coding47.9RatedGlobal rank #294/4 Measured dimensions
#7Multimodal45.7Estimated2/4 Measured dimensions
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Released
- Jul 2026
- Tokenizer
- Other
- Architecture
- text+image+audio->text
- Moderated
- No
- Supported parameters
- frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstoptemperaturetool_choicetoolstop_ktop_p
Rankings
Excels at
It's decent at
Falls behind in
Detailed scores
Updated: Sep 12, 2026Overall
undefined metric} other undefined metrics}}
Overall score60.0#48 / 112
Overall
undefined metric} other undefined metrics}}
Speed & latency
undefined metric} other undefined metrics}}
Output speed83.3 tok/s#52 / 81Time to first token2.63 s#56 / 81
Speed & latency
undefined metric} other undefined metrics}}
Pricing
undefined metric} other undefined metrics}}
Input price$1.00/M#113 / 186Output price$4.05/M#114 / 186
Pricing
undefined metric} other undefined metrics}}
Agents
V3.0undefined metric} other undefined metrics}} · Rated
Score48.6#4280% interval40.3–56.83/4 Measured dimensionsAgentic score57.6#46 / 77Terminal-Bench 2.063.8#27 / 53
Agents
V3.0undefined metric} other undefined metrics}} · Rated
Coding
V3.0undefined metric} other undefined metrics}} · Rated
Score47.9#2980% interval41.0–54.74/4 Measured dimensionsSciCode47.0%#41 / 89Coding score49.0#58 / 87
Coding
V3.0undefined metric} other undefined metrics}} · Rated
Reasoning
V3.0undefined metric} other undefined metrics}} · Estimated
Score55.280% interval44.4–66.02/4 Measured dimensionsGPQA87.2%#53 / 218HLE31.9%#52 / 216
Reasoning
V3.0undefined metric} other undefined metrics}} · Estimated
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Score51.380% interval37.3–65.31/4 Measured dimensionsKnowledge score65.1#52 / 83GPQA-D87.9#24 / 32
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Math
V3.0undefined metric} other undefined metrics}} · Provisional
Score65.480% interval49.0–81.71/4 Measured dimensionsMath score78.9#10 / 63AIME2697.1#2 / 14
Math
V3.0undefined metric} other undefined metrics}} · Provisional
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
Multimodal
V3.0undefined metric} other undefined metrics}} · Estimated
Score45.780% interval33.8–57.62/4 Measured dimensionsMultimodal Grounded score49.2#49 / 56MMMU-Pro73.5#28 / 31
Multimodal
V3.0undefined metric} other undefined metrics}} · Estimated
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
Score56.480% interval40.4–72.41/4 Measured dimensionsInstruction Following score85.2#32 / 52IFBench79.8#4 / 14
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
OpenRouter endpoints
3 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
BaseTen baseten/fp8 | $1/M | $4.05/M | 99.8% | — | — | undefined tokens / undefined tokens |
Together together | $1/M | $4.05/M | 98.5% | — | — | undefined tokens / undefined tokens |
DeepInfra deepinfra/fp8 | $0.950/M | $4.05/M | 93.3% | — | — | undefined tokens / undefined tokens |
Pricing Comparison
Compare Inkling API pricing across 20 providers. Prices range from $0.00000001/request to $10273.97/M. Future Hub offers the lowest rate at $0.00000001/request.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% | thinkingmachines/inkling:free | default | $75.00/M | $75.00/M | — | — | — | |
L1 99% | thinkingmachines/inkling:free | free | $0.00005/request | - | — | — | — | |
L1 100% | thinkingmachines/inkling | default | $75.00/M | $75.00/M | — | — | — | |
L1 99% | thinkingmachines/inkling:free | default | $547.50/M | $547.50/M | — | — | — | |
L1 70% | thinkingmachines/inkling | default | $75.00/M | $75.00/M | — | — | — | |
L1 47% | accounts/fireworks/models/inkling | 测试专用 | $10273.97/M | $10273.97/M | — | — | — | |
L1 100% | thinkingmachines/inkling:free | default | $75.00/M | $75.00/M | — | — | — | |
L1 100% | laohuang/thinkingmachines/inkling | default | $0.020/request | - | — | — | — | |
L1 0% | inkling | openrouter | $0.00000001/request | - | — | — | — | |
L1 0% L2 100% | inkling | price | -100%$0.0010/M Cache read$0.0002/MCache write$0.0010/MCache write 1h$0.0016/M | -100%$0.0040/M | — | — | — | |
L1 0% | inkling | 无限制 | $7.30/M Cache read$1.24/M | $29.57/M | — | — | — | |
L1 0% | inkling | model | $7.30/M Cache read$1.24/M | $29.57/M | — | — | — |
Alternatives & Similar Models
DeepSeek V4 Flash
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
DeepSeek V4 Pro
deepseek-v4-pro
DeepSeek V4 Pro is the professional-tier DeepSeek V4 model, targeting frontier reasoning, coding, and agent workflows with maximum capability.
MiniMax M3
minimax-m3
MiniMax M3 is MiniMax next-generation large language model, designed for advanced reasoning, long-context understanding, and high-quality multilingual dialogue.
GLM-5.1
glm-5-1
Zhipu GLM-5.1 is a next-generation GLM model aimed at frontier reasoning, coding, and bilingual agent applications.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
MiniMax M2.7
minimax-m2-7
MiniMax M2.7 is a high-tier M2-series model tuned for complex reasoning, long-context dialogue, and production-grade API workloads.
Frequently Asked Questions
- What benchmark data does Inkling include?
- LMSpeed shows Inkling benchmark context, API price, output speed, first-token latency, and provider data across 20 providers when those signals are available.
- What is the Inkling API price?
- Inkling has pricing from undefined provider} other undefined providers}}, ranging from $0.00000001/request to $10273.97/M. Future Hub has the lowest listed price.
- What does the Inkling API pricing table include?
- The Inkling API pricing table compares 20 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Inkling API pricing?
- Future Hub currently has the lowest listed Inkling price at $0.00000001/request across undefined provider} other undefined providers}}.
- Can I compare Inkling API price and speed together?
- Yes. LMSpeed shows Inkling API price, output speed, first-token latency, and provider health on the same page so you can compare cost and performance together.
- Is Inkling API free?
- Inkling does not currently have a free API tier on LMSpeed. All 20 providers charge per token.
