Llama 4 Maverick
Meta · Released 5 Apr 2025 · meta-llama/Llama-4-Maverick-17B-128E-Instruct
$0.24
Input price / 1M tokens
#7 of 61, cheapest first
$0.97
Output price / 1M tokens
#7 of 61, cheapest first
1M
Context window
#29 of 71
80
Output tokens / second
#40 of 61
69.8%
GPQA Diamond
#25 of 27
43.4%
LiveCodeBench
#5 of 7
59.6%
MMMU-Pro
#10 of 12
Summary
Llama 4 Maverick is an open-weight model from Meta, released on 5 Apr 2025.
At $0.24 input and $0.97 output per million tokens, it is #7 of 61 on input price, cheapest first.
Its 1M-token context window ranks #29 of 71.
Measured output speed is 80 tokens per second, #40 of 61.
Open-weight, natively multimodal mixture-of-experts model (17B active, 400B total, 128 experts) with a 1M-token context window. Meta positions it for multilingual assistant chat, image understanding and visual reasoning.
Benchmarks · 9 reported
| Benchmark | Score | Bar | Rank | Source |
|---|---|---|---|---|
| MMLU-Proinstruction-tuned, 0-shot, macro_avg/acc | 80.5% | – | Meta Llama 4 model card (Hugging Face)Self-reported | |
| GPQA Diamondinstruction-tuned, 0-shot | 69.8% | #25 of 27 | Meta Llama 4 model card (Hugging Face)Self-reported | |
| LiveCodeBenchinstruction-tuned, 0-shot, pass@1, problems 2024-10-01 to 2025-02-01 | 43.4% | #5 of 7 | Meta Llama 4 model card (Hugging Face)Self-reported | |
| MMMUinstruction-tuned, 0-shot | 73.4% | – | Meta Llama 4 model card (Hugging Face)Self-reported | |
| MMMU-Proinstruction-tuned, 0-shot, average of Standard and Vision | 59.6% | #10 of 12 | Meta Llama 4 model card (Hugging Face)Self-reported | |
| MathVistainstruction-tuned, 0-shot | 73.7% | – | Meta Llama 4 model card (Hugging Face)Self-reported | |
| ChartQAinstruction-tuned, 0-shot, relaxed accuracy | 90% | – | Meta Llama 4 model card (Hugging Face)Self-reported | |
| DocVQA (test)instruction-tuned, 0-shot, ANLS | 94.4% | – | Meta Llama 4 model card (Hugging Face)Self-reported | |
| MGSMinstruction-tuned, 0-shot, average/em | 92.3% | – | Meta Llama 4 model card (Hugging Face)Self-reported |
Self-reported means the lab ran the test itself. A rank appears only where several labs report the same version of a benchmark.
Similar models
Models with the most benchmarks in common with Llama 4 Maverick, and the closest scores.
GPQA Diamond
#25 of 27Graduate-level science questions
LiveCodeBench
#5 of 7Fresh competitive programming problems
MMMU-Pro
#10 of 12Hard questions that need an image
Compare Llama 4 Maverick side by side
Pre-filled with the two closest models. Swap any of them.
Sources
- license, context window, modalities, knowledge cutoff, release date, parameters, benchmarks: huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct, checked 2 Oct 2026
- Llama 4 is the latest Llama family listed by Meta: llama.com/llama/, checked 2 Oct 2026
- price (Amazon Bedrock on-demand and batch, us-east-1): aws.amazon.com/bedrock/pricing/, checked 2 Oct 2026
- Bedrock model ID, 1M context, 8K max output on Bedrock: docs.aws.amazon.com/bedrock/latest/userguide/model-card-meta-llama-4-maverick-17, checked 2 Oct 2026
- Groq retirement of Llama 4 Maverick: console.groq.com/docs/deprecations, checked 2 Oct 2026
- Together serverless list no longer includes Llama 4: together.ai/pricing, checked 2 Oct 2026
- output speed: artificialanalysis.ai/models/llama-4-maverick, checked 2 Oct 2026
Building on Llama 4 Maverick?
Get the architecture right before the bill arrives.
We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.