Llama 4 Maverick API Benchmarks, Pricing & Provider Data
Compare Llama 4 Maverick with another model
Choose a model to open its comparison page.
Llama 4 Maverick benchmark, API pricing, and provider data cover undefined API provider} other undefined API providers}}, with prices starting at $0.010/request.
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forw...
Category Performance
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
- Coverage
- 6 / 8
- Methodology
- V3.0
#1Reasoning46.4RatedGlobal rank #543/4 Measured dimensions
#2Math46.1Estimated3/4 Measured dimensions
#3Instruction following44.4Provisional1/4 Measured dimensions
#4Coding42.1Provisional1/4 Measured dimensions
#5Knowledge40.9Provisional1/4 Measured dimensions
#6Agents33Estimated2/4 Measured dimensions
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Total parameters
- 17B
- Released
- Apr 2025
- Knowledge cutoff
- 2024-08-31
- Tokenizer
- Llama4
- Architecture
- text+image->text
- Moderated
- No
- Supported parameters
- frequency_penaltylogit_biaslogprobsmax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Rankings
Excels at
It's decent at
Falls behind in
Detailed scores
Updated: Sep 12, 2026Overall
undefined metric} other undefined metrics}}
Overall score43.0#104 / 112
Overall
undefined metric} other undefined metrics}}
Speed & latency
undefined metric} other undefined metrics}}
Output speed96.3 tok/s#42 / 81Time to first token0.57 s#13 / 81
Speed & latency
undefined metric} other undefined metrics}}
Pricing
undefined metric} other undefined metrics}}
Input price$0.260/M#51 / 186Output price$0.910/M#45 / 186
Pricing
undefined metric} other undefined metrics}}
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Score3380% interval21.1–44.82/4 Measured dimensionsAA Agentic Index1.2#67 / 70Τ²-bench results17.8#79 / 82
Agents
V3.0undefined metric} other undefined metrics}} · Estimated
Coding
V3.0undefined metric} other undefined metrics}} · Provisional
Score42.180% interval28.2–56.01/4 Measured dimensionsLiveCodeBench39.7%#76 / 115SciCode31.7%#82 / 89
Coding
V3.0undefined metric} other undefined metrics}} · Provisional
Reasoning
V3.0undefined metric} other undefined metrics}} · Rated
Score46.4#5480% interval37.8–55.13/4 Measured dimensionsMMLU-Pro80.9%#57 / 129GPQA67.1%#150 / 218
Reasoning
V3.0undefined metric} other undefined metrics}} · Rated
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Score40.980% interval24.9–56.91/4 Measured dimensionsArtificial Analysis Intelligence Index14.5#100 / 117AA-GPQA Diamond67.1#95 / 113
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Math
V3.0undefined metric} other undefined metrics}} · Estimated
Score46.180% interval36.8–55.43/4 Measured dimensionsAIME39.0%#34 / 68MATH-50088.9%#36 / 73
Math
V3.0undefined metric} other undefined metrics}} · Estimated
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
80% interval30.8–69.20/4 Measured dimensionsAA-MMMU-Pro62.1#57 / 68Design Arena Website882.0#75 / 78
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
Score44.480% interval28.4–60.41/4 Measured dimensionsAA-IFBench43.0#62 / 84
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
OpenRouter endpoints
5 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
Parasail parasail/fp8 | $0.350/M | $1/M | 99.9% | — | — | undefined tokens / undefined tokens |
DigitalOcean digitalocean | $0.200/M | $0.696/M | 99.9% | — | — | undefined tokens / undefined tokens |
DeepInfra deepinfra/base | $0.200/M | $0.800/M | 99.8% | — | — | undefined tokens / undefined tokens |
Novita novita/fp8 | $0.270/M | $0.850/M | 99.7% | — | — | undefined tokens / undefined tokens |
Google google-vertex/us-east5 | $0.350/M | $1.15/M | — | — | — | undefined tokens / undefined tokens |
Pricing Comparison
Compare Llama 4 Maverick API pricing across 7 providers. Prices range from $0.010/request to $75.00/M. 素墨API offers the lowest rate at $0.010/request.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 99% | meta-llama/llama-4-maverick | default | $1.00/M | $4.00/M | — | — | — | |
L1 70% | llama-4-maverick | default | $0.010/request | - | — | — | — | |
L1 0% | meta-llama/llama-4-maverick | default | -26%$0.193/M | -15%$0.771/M | — | — | — | |
L1 0% | meta-llama/llama-4-maverick | default | $75.00/M | $75.00/M | — | — | — | |
L1 0% | llama-4-maverick-03-26-experimental | default | $75.00/M | $75.00/M | — | — | — | |
L1 0% | llama-4-maverick | nvidia | $1.00/request | - | — | — | — |
Alternatives & Similar Models
GPT-5.4
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
GPT-5.3 Codex
gpt-5-3-codex
OpenAI GPT-5.3 Codex is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks.
Claude Sonnet 4.6
claude-sonnet-4-6
Anthropic Claude Sonnet 4.6 extends the Sonnet line with improved tool use, coding reliability, and long-context performance for everyday production workloads.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
DeepSeek V3.2
deepseek-v3-2
DeepSeek V3.2 is an upgraded V3-series MoE model with stronger reasoning, coding, and math performance, widely available through OpenAI-compatible API relays.
GLM-4.7
glm-4-7
Zhipu GLM-4.7 is a flagship GLM release from Zhipu AI with advanced Chinese-English reasoning, coding, and agent features.
Frequently Asked Questions
- What benchmark data does Llama 4 Maverick include?
- LMSpeed shows Llama 4 Maverick benchmark context, API price, output speed, first-token latency, and provider data across 7 providers when those signals are available.
- What is the Llama 4 Maverick API price?
- Llama 4 Maverick has pricing from undefined provider} other undefined providers}}, ranging from $0.010/request to $75.00/M. 素墨API has the lowest listed price.
- What does the Llama 4 Maverick API pricing table include?
- The Llama 4 Maverick API pricing table compares 7 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Llama 4 Maverick API pricing?
- 素墨API currently has the lowest listed Llama 4 Maverick price at $0.010/request across undefined provider} other undefined providers}}.
- Is Llama 4 Maverick API free?
- Llama 4 Maverick does not currently have a free API tier on LMSpeed. All 7 providers charge per token.
