Meta
·Released on Dec 6, 2024

Llama 3.3 70B Instruct API Benchmarks, Pricing & Provider Data

Compare Llama 3.3 70B Instruct with another model

Choose a model to open its comparison page.

Share on X
LLM

Llama 3.3 70B Instruct API pricing covers undefined API provider} other undefined API providers}}, from $0.030/M to $57.53/M.

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

Cost
$0.030/ 1M · 8:1 in:out
$0.0047 in · $0.025 out

Specifications

Input and output token limits for this model, plus how it ranks on long-context understanding.

INPUT
131.1Ktokens
157.3 pages of text
OUTPUT
16.4Ktokens
8K128K1M4M
131.1K

Features

Technical Details

Input
Output
Total parameters
70B
Released
Dec 2024
Knowledge cutoff
2023-12-31
Tokenizer
Llama3
Architecture
text->text
Instruct type
llama3
Moderated
No
Supported parameters
frequency_penaltylogit_biaslogprobsmax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

OpenRouter endpoints

13 endpoints

Third-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.

Provider endpointInputOutput1d uptime30m latency30m throughputContext / output
Crusoe
crusoe/bf16
$0.250/M$0.750/M99.9%undefined tokens / undefined tokens
Google
google-vertex/us-central1
$0.720/M$0.720/M99.9%undefined tokens / undefined tokens
Groq
groq
$0.590/M$0.790/M99.8%undefined tokens / undefined tokens
Parasail
parasail/fp8
$0.220/M$0.500/M99.6%undefined tokens / undefined tokens
CoreWeave
coreweave/fp16
$0.710/M$0.710/M99.5%undefined tokens / undefined tokens
SambaNova
sambanova-turbo
$0.450/M$0.900/M99.4%undefined tokens / undefined tokens
Google
google-vertex
$0.720/M$0.720/M99.4%undefined tokens / undefined tokens
Together
together
$1.04/M$1.04/M99.0%undefined tokens / undefined tokens
DeepInfra
deepinfra/turbo
$0.100/M$0.320/M98.6%undefined tokens / undefined tokens
Cloudflare
cloudflare/fp8
$0.293/M$2.25/M98.6%undefined tokens / undefined tokens
Novita
novita/bf16
$0.135/M$0.400/M98.4%undefined tokens / undefined tokens
AkashML
akashml/fp8
$0.200/M$0.520/M97.9%undefined tokens / undefined tokens
Nebius
nebius/fp8
$0.130/M$0.400/M40.9%undefined tokens / undefined tokens

Pricing Comparison

Compare Llama 3.3 70B Instruct API pricing across 6 providers. Prices range from $0.030/M to $57.53/M. 云AI offers the lowest rate at $0.030/M.

ProviderHealthModel VariantGroupInput ($/M)Output ($/M)Speed (t/s)First tokenAudit
L1
100%
llama-3.3-70b-instruct
enterprise_emergency_standby
$57.53/M
$57.53/M
L1
100%
llama-3.3-70b-instruct
azure03-0
$0.030/M
$0.030/M
L1
100%
llama-3.3-70b-instruct
Self-Deployed-3
$0.160/M
$0.160/M
L1
100%
llama-3.3-70b-instruct
Self-Deployed-3
$0.160/M
$0.160/M
L1
100%
llama-3.3-70b-instruct
纯AZ
$0.074/M
$0.074/M

Alternatives & Similar Models

Frequently Asked Questions

What benchmark data does Llama 3.3 70B Instruct include?
LMSpeed shows Llama 3.3 70B Instruct benchmark context, API price, output speed, first-token latency, and provider data across 6 providers when those signals are available.
What is the Llama 3.3 70B Instruct API price?
Llama 3.3 70B Instruct has pricing from undefined provider} other undefined providers}}, ranging from $0.030/M to $57.53/M. 云AI has the lowest listed price.
What does the Llama 3.3 70B Instruct API pricing table include?
The Llama 3.3 70B Instruct API pricing table compares 6 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
Which provider has the cheapest Llama 3.3 70B Instruct API pricing?
云AI currently has the lowest listed Llama 3.3 70B Instruct price at $0.030/M across undefined provider} other undefined providers}}.
Is Llama 3.3 70B Instruct API free?
Llama 3.3 70B Instruct does not currently have a free API tier on LMSpeed. All 6 providers charge per token.

Also known as

llama-3.3-70b-instructmeta-llama/Llama-3.3-70B-Instructmeta-llama/llama-3.3-70b-instructmeta-llama/llama-3.3-70b-instruct:freemeta/llama-3.3-70b-instruct

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.Standard benchmark data may include BenchLM and other public sources.

Build with LMSpeed Data

Free

Free public API for LLM pricing, benchmarks & provider data

  • Real-time pricing & availability
  • Speed & latency benchmarks
  • 300+ models, 600+ providers
  • 1,000 requests/day
View API Documentation