Llama 3.2 Nemoretriever 300m Embed v1 API Benchmarks, Pricing & Provider Data
Compare Llama 3.2 Nemoretriever 300m Embed v1 with another model
Choose a model to open its comparison page.
Llama 3.2 Nemoretriever 300m Embed v1 API pricing covers 3 API providers, from $0.010/request to $75.00/M.
Meta Llama 3.2 Nemoretriever 300m Embed v1 is an embedding model, designed for generating vector representations of text for retrieval and semantic search.
Specifications
Pricing Comparison
Compare Llama 3.2 Nemoretriever 300m Embed v1 API pricing across 3 providers. Prices range from $0.010/request to $75.00/M. 素墨API offers the lowest rate at $0.010/request.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% | llama-3.2-nemoretriever-300m-embed-v1 | default | $0.010/request | - | — | — | — | |
L1 100% | nvidia/llama-3.2-nemoretriever-300m-embed-v1 | default | $0.010/request | - | — | — | — | |
L1 99% | nvidia/llama-3.2-nemoretriever-300m-embed-v1 | default | $75.00/M | $75.00/M | — | — | — |
Alternatives & Similar Models
GLM-5.1
glm-5-1
Zhipu GLM-5.1 is a next-generation GLM model aimed at frontier reasoning, coding, and bilingual agent applications.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
DeepSeek V3.2
deepseek-v3-2
DeepSeek V3.2 is an upgraded V3-series MoE model with stronger reasoning, coding, and math performance, widely available through OpenAI-compatible API relays.
MiniMax M2.5
minimax-m2-5
MiniMax M2.5 is MiniMax's flagship text model for coding and agents, with SOTA-level programming and agentic performance, improved token efficiency, and fast high-TPS API deployment.
MiniMax M2.7
minimax-m2-7
MiniMax M2.7 is a high-tier M2-series model tuned for complex reasoning, long-context dialogue, and production-grade API workloads.
GPT-OSS
gpt-oss
GPT-OSS is an open-weight language model family designed for self-hosted inference, research, and cost-efficient alternatives to proprietary GPT-class models.
Frequently Asked Questions
- What benchmark data does Llama 3.2 Nemoretriever 300m Embed v1 include?
- LMSpeed shows Llama 3.2 Nemoretriever 300m Embed v1 benchmark context, API price, output speed, first-token latency, and provider data across 3 providers when those signals are available.
- What is the Llama 3.2 Nemoretriever 300m Embed v1 API price?
- Llama 3.2 Nemoretriever 300m Embed v1 has pricing from 3 providers, ranging from $0.010/request to $75.00/M. 素墨API has the lowest listed price.
- What does the Llama 3.2 Nemoretriever 300m Embed v1 API pricing table include?
- The Llama 3.2 Nemoretriever 300m Embed v1 API pricing table compares 3 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Llama 3.2 Nemoretriever 300m Embed v1 API pricing?
- 素墨API currently has the lowest listed Llama 3.2 Nemoretriever 300m Embed v1 price at $0.010/request across 3 providers.
- Is Llama 3.2 Nemoretriever 300m Embed v1 API free?
- Llama 3.2 Nemoretriever 300m Embed v1 does not currently have a free API tier on LMSpeed. All 3 providers charge per token.
