Choose a model to open its comparison page.
Nemotron Nano 9B V2 API pricing covers 5 API providers, from $0.0010/request to $0.100/request. Nemotron Nano 9B V2 free API options are available from 1 provider.
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and.....
Input and output token limits for this model, plus how it ranks on long-context understanding.
Compare Nemotron Nano 9B V2 API pricing across 4 providers. Prices range from $0.0010/request to $0.100/request. CM-API 公益站 offers the lowest rate at $0.0010/request. 1 provider offers free API credits or a free tier.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% | nvidia/nemotron-nano-9b-v2:free | free | $0.0010/request | - | — | — | — | |
Dext API Free | L1 100% | nemotron-nano-9b-v2 | 公益 | Free | Free | — | — | — |
L1 96% | nemotron-nano-9b-v2 | 小叶币 | $0.050/request | - | — | — | — | |
L1 0% | nvidia/nemotron-nano-9b-v2:free | default | $0.010/request | - | — | — | — |
gpt-oss
GPT-OSS is an open-weight language model family designed for self-hosted inference, research, and cost-efficient alternatives to proprietary GPT-class models.
glm-4-5-air
Zhipu AI GLM-4.5 Air is a lightweight GLM-4.5 variant optimized for low-latency Chinese and English dialogue, retrieval-augmented apps, and edge deployments.
lfm-2-5-1-2b-instruct
LFM 2.5 1.2B Instruct is a compact language model in the LFM series, optimized for low-latency responses and efficient inference.
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
claude-haiku-4-5
Anthropic Claude Haiku 4.5 delivers fast, low-cost responses while retaining solid instruction following for chat, classification, and lightweight coding.
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.