Skip to content

Gemma 4 26B A4B

Google · Released 31 Mar 2026 · gemma-4-26b-a4b-it

Compare this model
Open weightBudgetReasoning

–

Input price / 1M tokens

–

Output price / 1M tokens

262K

Context window

#54 of 71

–

Output tokens / second

77.1%

LiveCodeBench

#3 of 7

82.3%

GPQA Diamond

#22 of 27

68.2%

τ²-bench

#6 of 6

8.7%

Humanity’s Last Exam (no tools)

#22 of 23

Summary

Gemma 4 26B A4B is an open-weight model from Google, released on 31 Mar 2026.

Its 262K-token context window ranks #54 of 71.

Mixture-of-experts Gemma 4 open model: 25.2B total and 3.8B active parameters, 256K context. It runs almost as fast as a 4B model.

Benchmarks · 9 reported

BenchmarkScoreBarRankSource
MMLU Proinstruction-tuned82.6%–Gemma 4 model cardSelf-reported
AIME 2026no tools88.3%–Gemma 4 model cardSelf-reported
LiveCodeBench77.1%#3 of 7Gemma 4 model cardSelf-reported
Codeforces1,718Elo-style–Gemma 4 model cardSelf-reported
GPQA Diamond82.3%#22 of 27Gemma 4 model cardSelf-reported
τ²-benchaverage over 368.2%#6 of 6Gemma 4 model cardSelf-reported
Humanity’s Last Exam (no tools)no tools (17.2% with search)8.7%#22 of 23Gemma 4 model cardSelf-reported
MMMU-Pro73.8%#6 of 12Gemma 4 model cardSelf-reported
MRCR v2 (8-needle)128k average44.1%#12 of 13Gemma 4 model cardSelf-reported

Self-reported means the lab ran the test itself. A rank appears only where several labs report the same version of a benchmark.

Similar models

Models with the most benchmarks in common with Gemma 4 26B A4B, and the closest scores.

Compare 4 side by side

LiveCodeBench

#3 of 7

Fresh competitive programming problems

GPQA Diamond

#22 of 27

Graduate-level science questions

τ²-bench

#6 of 6

Tool use with a simulated customer

Humanity’s Last Exam (no tools)

#22 of 23

Expert questions, no search or code

Compare Gemma 4 26B A4B side by side

Pre-filled with the two closest models. Swap any of them.

Open comparison

Sources

Building on Gemma 4 26B A4B?

Get the architecture right before the bill arrives.

We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.