Llama 3.3 API Benchmarks, Pricing & Provider Data
Compare Llama 3.3 with another model
Choose a model to open its comparison page.
Llama 3.3 API pricing covers 37 API providers, from $0.0027/M to $75.00/M. The page also shows measured API speed and first-token latency.
Meta Llama 3.3 is an updated Llama 3 open model with improved instruction following, multilingual support, and efficient inference.
Specifications
Pricing Comparison
Compare Llama 3.3 API pricing across 37 providers. Prices range from $0.0027/M to $75.00/M. 云AI offers the lowest rate at $0.0027/M.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% | llama-3.3-70b | default | $0.0068/M | $0.020/M | — | — | — | |
L1 100% | llama-3.3-70b | mix | $0.0027/M | $0.0081/M | — | — | — | |
L1 100% | llama-3.3-70b | default | $0.014/M | $0.041/M | — | — | — | |
L1 100% | llama-3.3-70b | default | $0.0068/M | $0.020/M | — | — | — | |
L1 99% | llama-3.3-70b | default | $0.011/M | $0.032/M | — | — | — | |
L1 100% | llama-3.3-70b | default | $0.013/M | $0.038/M | — | — | — | |
L1 99% | llama-3.3-70b | default | $0.013/M | $0.038/M | — | — | — | |
L1 100% | llama-3.3-70b | default | $0.0068/M | $0.020/M | — | — | — | |
L1 0% | llama-3.3-70b | default | $75.00/M | $75.00/M | — | — | — |
Alternatives & Similar Models
Gemini 2.5 Pro
gemini-2-5-pro
Google Gemini 2.5 Pro is Google advanced multimodal model with a 1M-token context window, strong STEM reasoning, and native support for images, audio, and video understanding.
Gemini 2.5 Flash
gemini-2-5-flash
Google Gemini 2.5 Flash is a fast multimodal model balancing speed and intelligence for chat, tool use, and large-context workloads at lower cost than Pro tiers.
DeepSeek V3
deepseek-v3
DeepSeek V3 is DeepSeek flagship MoE language model with 671B total parameters, delivering strong performance in reasoning, coding, and multilingual tasks at competitive inference cost.
GLM-4.7
glm-4-7
Zhipu GLM-4.7 is a flagship GLM release from Zhipu AI with advanced Chinese-English reasoning, coding, and agent features.
Qwen3
qwen3
Alibaba Qwen3 is the Qwen family's flagship LLM series with dense and MoE variants, seamless thinking/non-thinking modes, and leading open-source performance in math, code, and agent tasks.
GPT-5.4
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
Frequently Asked Questions
- What benchmark data does Llama 3.3 include?
- LMSpeed shows Llama 3.3 benchmark context, API price, output speed, first-token latency, and provider data across 37 providers when those signals are available.
- What is the Llama 3.3 API price?
- Llama 3.3 has pricing from 37 providers, ranging from $0.0027/M to $75.00/M. 云AI has the lowest listed price.
- What does the Llama 3.3 API pricing table include?
- The Llama 3.3 API pricing table compares 37 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Llama 3.3 API pricing?
- 云AI currently has the lowest listed Llama 3.3 price at $0.0027/M across 37 providers.
- Can I compare Llama 3.3 API price and speed together?
- Yes. LMSpeed shows Llama 3.3 API price, output speed, first-token latency, and provider health on the same page so you can compare cost and performance together.
- Is Llama 3.3 API free?
- Llama 3.3 does not currently have a free API tier on LMSpeed. All 37 providers charge per token.
