Llama 4 Scout
Meta · Released 5 Apr 2025 · meta-llama/Llama-4-Scout-17B-16E-Instruct
$0.17
Input price / 1M tokens
#5 of 61, cheapest first
$0.66
Output price / 1M tokens
#5 of 61, cheapest first
10M
Context window
#1 of 71
108
Output tokens / second
#27 of 61
57.2%
GPQA Diamond
#27 of 27
32.8%
LiveCodeBench
#7 of 7
52.2%
MMMU-Pro
#11 of 12
Summary
Llama 4 Scout is an open-weight model from Meta, released on 5 Apr 2025.
At $0.17 input and $0.66 output per million tokens, it is #5 of 61 on input price, cheapest first.
Its 10M-token context window ranks #1 of 71.
Measured output speed is 108 tokens per second, #27 of 61.
Open-weight, natively multimodal mixture-of-experts model (17B active, 109B total, 16 experts) with a 10M-token context window. Meta positions it for assistant chat, visual reasoning and long-document work.
Benchmarks · 9 reported
| Benchmark | Score | Bar | Rank | Source |
|---|---|---|---|---|
| MMLU-Proinstruction-tuned, 0-shot, macro_avg/acc | 74.3% | – | Meta Llama 4 model card (Hugging Face)Self-reported | |
| GPQA Diamondinstruction-tuned, 0-shot | 57.2% | #27 of 27 | Meta Llama 4 model card (Hugging Face)Self-reported | |
| LiveCodeBenchinstruction-tuned, 0-shot, pass@1, problems 2024-10-01 to 2025-02-01 | 32.8% | #7 of 7 | Meta Llama 4 model card (Hugging Face)Self-reported | |
| MMMUinstruction-tuned, 0-shot | 69.4% | – | Meta Llama 4 model card (Hugging Face)Self-reported | |
| MMMU-Proinstruction-tuned, 0-shot, average of Standard and Vision | 52.2% | #11 of 12 | Meta Llama 4 model card (Hugging Face)Self-reported | |
| MathVistainstruction-tuned, 0-shot | 70.7% | – | Meta Llama 4 model card (Hugging Face)Self-reported | |
| ChartQAinstruction-tuned, 0-shot, relaxed accuracy | 88.8% | – | Meta Llama 4 model card (Hugging Face)Self-reported | |
| DocVQA (test)instruction-tuned, 0-shot, ANLS | 94.4% | – | Meta Llama 4 model card (Hugging Face)Self-reported | |
| MGSMinstruction-tuned, 0-shot, average/em | 90.6% | – | Meta Llama 4 model card (Hugging Face)Self-reported |
Self-reported means the lab ran the test itself. A rank appears only where several labs report the same version of a benchmark.
Similar models
Models with the most benchmarks in common with Llama 4 Scout, and the closest scores.
GPQA Diamond
#27 of 27Graduate-level science questions
LiveCodeBench
#7 of 7Fresh competitive programming problems
MMMU-Pro
#11 of 12Hard questions that need an image
Compare Llama 4 Scout side by side
Pre-filled with the two closest models. Swap any of them.
Sources
- license, context window, modalities, knowledge cutoff, release date, parameters, benchmarks: huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct, checked 2 Oct 2026
- Llama 4 is the latest Llama family listed by Meta: llama.com/llama/, checked 2 Oct 2026
- price (Amazon Bedrock on-demand and batch, us-east-1): aws.amazon.com/bedrock/pricing/, checked 2 Oct 2026
- Bedrock model ID, 10M context, 8K max output on Bedrock: docs.aws.amazon.com/bedrock/latest/userguide/model-card-meta-llama-4-scout-17b-i, checked 2 Oct 2026
- Groq retirement of Llama 4 Scout: console.groq.com/docs/deprecations, checked 2 Oct 2026
- Together serverless list no longer includes Llama 4: together.ai/pricing, checked 2 Oct 2026
- output speed: artificialanalysis.ai/models/llama-4-scout, checked 2 Oct 2026
Building on Llama 4 Scout?
Get the architecture right before the bill arrives.
We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.