Llama 4 Scout API Benchmarks, Pricing & Provider Data
Compare Llama 4 Scout with another model
Choose a model to open its comparison page.
Llama 4 Scout benchmark, API pricing, and provider data cover 24 API providers, with prices starting at $0.0077/M. Llama 4 Scout free API options are available from 1 provider. The page also shows measured API speed and first-token latency.
Meta Llama 4 Scout is a compact open-weight model in the Llama 4 family, designed for efficient inference, on-device deployment, and low-latency agent workloads.
Category Performance
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
- Coverage
- 6 / 8
- Methodology
- V3.0
#1Math44.4Estimated3/4 Measured dimensions
#2Instruction following43.2Provisional1/4 Measured dimensions
#3Reasoning41.8RatedGlobal rank #603/4 Measured dimensions
#4Knowledge38.4Provisional1/4 Measured dimensions
#5Coding34.3Provisional1/4 Measured dimensions
#6Agents32.7Estimated2/4 Measured dimensions
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Total parameters
- 17B
- Released
- Apr 2025
- Knowledge cutoff
- 2024-08-31
- Tokenizer
- Llama4
- Architecture
- text+image->text
- Moderated
- No
- Supported parameters
- frequency_penaltylogit_biasmax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Rankings
Excels at
It's decent at
Falls behind in
Detailed scores
Updated: Sep 1, 2026Overall
undefined metric} other undefined metrics}}
Overall score43.0#91 / 100
Overall
undefined metric} other undefined metrics}}
Speed & latency
undefined metric} other undefined metrics}}
Output speed133.0 tok/s#24 / 77Time to first token0.57 s#14 / 77
Speed & latency
undefined metric} other undefined metrics}}
Pricing
undefined metric} other undefined metrics}}
Input price$0.180/M#31 / 181Output price$0.660/M#31 / 181
Pricing
undefined metric} other undefined metrics}}
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Score32.780% interval20.9–44.52/4 Measured dimensionsAA Agentic Index1.1#64 / 65Τ²-bench results15.5#81 / 82
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Coding
V3.0undefined metric} other undefined metrics}} · Provisional
Score34.380% interval20.4–48.21/4 Measured dimensionsLiveCodeBench29.9%#89 / 115SciCode17.0%#201 / 206
Coding
V3.0undefined metric} other undefined metrics}} · Provisional
Reasoning
V3.0undefined metric} other undefined metrics}} · Rated
Score41.8#6080% interval33.2–50.53/4 Measured dimensionsMMLU-Pro75.2%#85 / 129GPQA58.7%#170 / 213
Reasoning
V3.0undefined metric} other undefined metrics}} · Rated
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Score38.480% interval22.4–54.41/4 Measured dimensionsArtificial Analysis Intelligence Index10.3#103 / 111AA-GPQA Diamond58.7#98 / 108
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Math
V3.0undefined metric} other undefined metrics}} · Estimated
Score44.480% interval35.1–53.63/4 Measured dimensionsAIME28.3%#42 / 68MATH-50084.4%#44 / 73
Math
V3.0undefined metric} other undefined metrics}} · Estimated
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
80% interval30.8–69.20/4 Measured dimensionsAA-MMMU-Pro52.9#61 / 65Design Arena Website761.0#75 / 75
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
Score43.280% interval27.2–59.21/4 Measured dimensionsAA-IFBench39.5#70 / 84
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
OpenRouter endpoints
4 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
Google google-vertex/us-east5 | $0.250/M | $0.700/M | 99.9% | — | — | undefined tokens / undefined tokens |
DeepInfra deepinfra/fp8 | $0.100/M | $0.300/M | 99.9% | — | — | undefined tokens / undefined tokens |
Novita novita/bf16 | $0.180/M | $0.590/M | 99.8% | — | — | undefined tokens / undefined tokens |
Groq groq | $0.110/M | $0.340/M | 98.4% | — | — | undefined tokens / undefined tokens |
Pricing Comparison
Compare Llama 4 Scout API pricing across 23 providers. Prices range from $0.0077/M to $77.05/M. 云AI offers the lowest rate at $0.0077/M. 1 provider offers free API credits or a free tier.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 100% | meta-llama/llama-4-scout | default | $3.75/M | $15.00/M | — | — | — | |
L1 100% | llama-4-scout | default | $0.300/M | $0.945/M | — | — | — | |
L1 99% | meta-llama/llama-4-scout | default | $0.490/M | $1.47/M | — | — | — | |
L1 99% | meta-llama/llama-4-scout-17b-16e-instruct | default | $10.27/M | $10.27/M | — | — | — | |
L1 100% | llama-4-scout | default | $0.010/request | - | — | — | — | |
L1 99% | llama-4-scout | default | -77%$0.041/M | -80%$0.129/M | — | — | — | |
L1 100% | llama-4-scout | mix | -96%$0.0077/M | -96%$0.024/M | — | — | — | |
L1 100% | llama-4-scout-17b-16e-instruct | Model | -89%$0.019/M | -91%$0.060/M | — | — | — | |
L1 100% | llama-4-scout-17b-16e-instruct | default | -54%$0.082/M | -61%$0.259/M | — | — | — | |
L1 100% | llama-4-scout | default | -54%$0.082/M | -61%$0.259/M | — | — | — | |
Dext API Free | L1 55% | llama-4-scout-17b-16e-instruct | 公益 | Free | Free | — | — | — |
L1 0% | llama-4-scout-17b-16e-instruct | 小叶币 | $0.050/request | - | — | — | — | |
L1 0% | llama-4-scout | default | -66%$0.062/M | -71%$0.194/M | — | — | — | |
L1 0% | llama-4-scout-17b | default | $75.00/M | $75.00/M | — | — | — | |
L1 0% | meta-llama/llama-4-scout | default | $75.00/M | $75.00/M | — | — | — | |
L1 0% | llama-4-scout | default | -77%$0.041/M | -80%$0.129/M | — | — | — | |
L1 0% | meta-llama/llama-4-scout | default | -35%$0.116/M | -34%$0.436/M | — | — | — |
Alternatives & Similar Models
Gemini 2.5 Pro
gemini-2-5-pro
Google Gemini 2.5 Pro is Google advanced multimodal model with a 1M-token context window, strong STEM reasoning, and native support for images, audio, and video understanding.
GPT-OSS
gpt-oss
GPT-OSS is an open-weight language model family designed for self-hosted inference, research, and cost-efficient alternatives to proprietary GPT-class models.
GPT-5.4
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
Claude Sonnet 4.6
claude-sonnet-4-6
Anthropic Claude Sonnet 4.6 extends the Sonnet line with improved tool use, coding reliability, and long-context performance for everyday production workloads.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
GLM-5
glm-5
Zhipu GLM-5 is Zhipu flagship GLM series model with enhanced reasoning, agent capabilities, and strong performance on Chinese enterprise and coding scenarios.
Frequently Asked Questions
- What benchmark data does Llama 4 Scout include?
- LMSpeed shows Llama 4 Scout benchmark context, API price, output speed, first-token latency, and provider data across 24 providers when those signals are available.
- What is the Llama 4 Scout API price?
- Llama 4 Scout has pricing from 24 providers, ranging from $0.0077/M to $77.05/M. 云AI has the lowest listed price.
- What does the Llama 4 Scout API pricing table include?
- The Llama 4 Scout API pricing table compares 24 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Llama 4 Scout API pricing?
- 云AI currently has the lowest listed Llama 4 Scout price at $0.0077/M across 24 providers.
- Can I compare Llama 4 Scout API price and speed together?
- Yes. LMSpeed shows Llama 4 Scout API price, output speed, first-token latency, and provider health on the same page so you can compare cost and performance together.
- Is Llama 4 Scout API free?
- Yes, Llama 4 Scout free API options are available through 1 providerundefined other undefined} on LMSpeed, including Dext API. These providers offer free API credits or a free tier with no per-token charges.
- Where can I get Llama 4 Scout free API access?
- LMSpeed currently lists 1 free API providerundefined other undefined} for Llama 4 Scout: Dext API. Check each provider row before using it because free tier limits can change.
