Llama 3.2 3B Instruct API Benchmarks, Pricing & Provider Data
Compare Llama 3.2 3B Instruct with another model
Choose a model to open its comparison page.
Llama 3.2 3B Instruct API pricing covers undefined API provider} other undefined API providers}}, from $0.010/request to $0.222/M. Llama 3.2 3B Instruct free API options are available from undefined provider} other undefined providers}}.
Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with ...
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Total parameters
- 3B
- Released
- Sep 2024
- Knowledge cutoff
- 2023-12-31
- Tokenizer
- Llama3
- Architecture
- text->text
- Instruct type
- llama3
- Moderated
- No
- Supported parameters
- frequency_penaltylogit_biaslogprobsmax_tokensmin_ppresence_penaltyrepetition_penaltyseedstopstructured_outputstemperaturetop_ktop_logprobstop_p
OpenRouter endpoints
2 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
Parasail parasail/bf16 | $0.050/M | $0.330/M | 100.0% | — | — | undefined tokens / undefined tokens |
Cloudflare cloudflare | $0.051/M | $0.335/M | 92.6% | — | — | undefined tokens / undefined tokens |
Pricing Comparison
Compare Llama 3.2 3B Instruct API pricing across 9 providers. Prices range from $0.010/request to $0.222/M. uglycat offers the lowest rate at $0.010/request. 1 provider offers free API credits or a free tier.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% | llama-3.2-3b-instruct | default | $0.027/M | $0.027/M | — | — | — | |
兔子API Free | L1 99% | llama-3.2-3b-instruct | default | Free | Free | — | — | — |
L1 100% | llama-3.2-3b-instruct | Self-Deployed-2 | $0.148/M | $0.074/M | — | — | — | |
L1 99% | llama-3.2-3b-instruct | Self-Deployed-2 | $0.148/M | $0.074/M | — | — | — | |
L1 100% | llama-3.2-3b-instruct | default | $0.068/M | $0.034/M | — | — | — | |
L1 0% | meta-llama/llama-3.2-3b-instruct:free | default | $0.010/request | - | — | — | — |
Alternatives & Similar Models
MiniMax M2.5
minimax-m2-5
MiniMax M2.5 is MiniMax's flagship text model for coding and agents, with SOTA-level programming and agentic performance, improved token efficiency, and fast high-TPS API deployment.
DeepSeek V4 Flash
deepseek-v4-flash
DeepSeek V4 Flash is a fast, cost-efficient language model in the DeepSeek V4 family, optimized for low-latency chat, coding assistance, and high-throughput API workloads while retaining strong reasoning quality.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
DeepSeek V3.2
deepseek-v3-2
DeepSeek V3.2 is an upgraded V3-series MoE model with stronger reasoning, coding, and math performance, widely available through OpenAI-compatible API relays.
MiniMax M2.7
minimax-m2-7
MiniMax M2.7 is a high-tier M2-series model tuned for complex reasoning, long-context dialogue, and production-grade API workloads.
GLM-4.5 Air
glm-4-5-air
Zhipu AI GLM-4.5 Air is a lightweight GLM-4.5 variant optimized for low-latency Chinese and English dialogue, retrieval-augmented apps, and edge deployments.
Frequently Asked Questions
- What benchmark data does Llama 3.2 3B Instruct include?
- LMSpeed shows Llama 3.2 3B Instruct benchmark context, API price, output speed, first-token latency, and provider data across 10 providers when those signals are available.
- What is the Llama 3.2 3B Instruct API price?
- Llama 3.2 3B Instruct has pricing from undefined provider} other undefined providers}}, ranging from $0.010/request to $0.222/M. uglycat has the lowest listed price.
- What does the Llama 3.2 3B Instruct API pricing table include?
- The Llama 3.2 3B Instruct API pricing table compares 10 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Llama 3.2 3B Instruct API pricing?
- uglycat currently has the lowest listed Llama 3.2 3B Instruct price at $0.010/request across undefined provider} other undefined providers}}.
- Is Llama 3.2 3B Instruct API free?
- Yes, Llama 3.2 3B Instruct free API options are available through 1 providerundefined other undefined} on LMSpeed, including 兔子API. These providers offer free API credits or a free tier with no per-token charges.
- Where can I get Llama 3.2 3B Instruct free API access?
- LMSpeed currently lists 1 free API providerundefined other undefined} for Llama 3.2 3B Instruct: 兔子API. Check each provider row before using it because free tier limits can change.
