Choose a model to open its comparison page.
Llama Nemotron Embed VL 1B V2 API pricing covers 9 API providers, from $0.0000001/request to $375.00/M. Llama Nemotron Embed VL 1B V2 free API options are available from 5 providers.
The Llama Nemotron Embed VL 1B V2 embedding model is optimized for multimodal question-answering retrieval. The model can embed 'documents' in the form of image, text, or image and text...
Input and output token limits for this model, plus how it ranks on long-context understanding.
Compare Llama Nemotron Embed VL 1B V2 API pricing across 4 providers. Prices range from $0.0000001/request to $375.00/M. 91VIP API offers the lowest rate at $0.0000001/request. 5 providers offer free API credits or a free tier.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% | nvidia/llama-nemotron-embed-vl-1b-v2 | default | $2.00/M | $2.00/M | — | — | — | |
Dext API Free | L1 100% | llama-nemotron-embed-vl-1b-v2 | 公益 | Free | Free | — | — | — |
L1 99% | nvidia/llama-nemotron-embed-vl-1b-v2 | default | $375.00/M | $375.00/M | — | — | — | |
91VIP API Free | L1 95% | nvidia/llama-nemotron-embed-vl-1b-v2 | default | Free | Free | — | — | — |
L1 95% | nvidia/llama-nemotron-embed-vl-1b-v2:free | default | $0.0000001/request | - | — | — | — | |
梦德 API Free | L1 75% | llama-nemotron-embed-vl-1b-v2 | default | Free | Free | — | — | — |
L1 75% | nvidia/llama-nemotron-embed-vl-1b-v2 | default | Free | Free | — | — | — | |
L1 75% | nvidia/llama-nemotron-embed-vl-1b-v2:free | default | Free | Free | — | — | — |
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
gpt-5-4-mini
OpenAI GPT-5.4 Mini is a compact language model in the GPT-5 series, optimized for quick responses and high throughput.
claude-opus-4-6
Anthropic Claude Opus 4.6 is the most capable Claude Opus tier, optimized for complex analysis, long-horizon coding, and high-stakes enterprise reasoning workloads.
claude-sonnet-4-6
Anthropic Claude Sonnet 4.6 extends the Sonnet line with improved tool use, coding reliability, and long-context performance for everyday production workloads.
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
deepseek-v4-pro
DeepSeek V4 Pro is the professional-tier DeepSeek V4 model, targeting frontier reasoning, coding, and agent workflows with maximum capability.