Skip to content

Gemini 3.8 Flash

Google · Released 2 Sep 2026 · gemini-3.8-flash

Compare this model
ProprietaryCodingAgentsReasoningMultimodal

$0.75

Input price / 1M tokens

#18 of 61, cheapest first

$3.75

Output price / 1M tokens

#21 of 61, cheapest first

1.05M

Context window

#11 of 71

249

Output tokens / second

#5 of 61

73.7%

DeepSWE v1.1

#5 of 17

89.4%

Terminal-Bench 2.1

#2 of 17

19.1%

Terminal-Bench 4.0

#9 of 10

59%

OSWorld 2.0

#7 of 10

Summary

Gemini 3.8 Flash is a proprietary model from Google, released on 2 Sep 2026.

At $0.75 input and $3.75 output per million tokens, it is #18 of 61 on input price, cheapest first.

Its 1.05M-token context window ranks #11 of 71.

Measured output speed is 249 tokens per second, #5 of 61.

Google's most intelligent Flash model, built for long-horizon software engineering, autonomous agents and complex enterprise workflows at Flash speed and cost.

Benchmarks · 13 reported

BenchmarkScoreBarRankSource
DeepSWE v1.1mini-swe agent harness, high thinking73.7%#5 of 17Google Gemini 3.8 Flash evaluation reportSelf-reported
Terminal-Bench 2.1Terminus 2 harness89.4%#2 of 17Google Gemini 3.8 Flash evaluation reportSelf-reported
Terminal-Bench 4.0from official public leaderboard, highest-scoring thinking level19.1%#9 of 10Terminal-Bench leaderboard, cited in Google's evaluation report
HLE-Verifiedfull 1,811-item verified set54.9%–Google Gemini 3.8 Flash evaluation reportSelf-reported
OSWorld 2.0partial score, batched tool calls, max over 3 runs, single attempt per run59%#7 of 10Google Gemini 3.8 Flash evaluation reportSelf-reported
GDPVal-AA v2Artificial Analysis leaderboard1,545Elo-style–Artificial Analysis, cited in Google's evaluation report
Vals Finance Agent v261.4%–Vals.ai, cited in Google's evaluation report
Harvey's Legal Agent Benchmarkall-pass rate10%–Vals.ai, cited in Google's evaluation report
GDP.PDFall-pass rate35%–Google Gemini 3.8 Flash evaluation reportSelf-reported
CharXiv Reasoningno tools86.2%#1 of 6Google Gemini 3.8 Flash evaluation reportSelf-reported
LVBenchstatic, no tools, 1024 frames (87.8% in agentic mode)87.1%–Google Gemini 3.8 Flash evaluation reportSelf-reported
BioMysteryBenchhuman-solvable set (56.5% on human-difficult set)88.8%–Google Gemini 3.8 Flash evaluation reportSelf-reported
LABBench2terminal with bioinformatics tools and internet; macro-average of 11 sub-tasks86.2%–Google Gemini 3.8 Flash evaluation reportSelf-reported

Self-reported means the lab ran the test itself. A rank appears only where several labs report the same version of a benchmark.

Similar models

Models with the most benchmarks in common with Gemini 3.8 Flash, and the closest scores.

Compare 4 side by side

DeepSWE v1.1

#5 of 17

Long coding tasks in real repos

Terminal-Bench 2.1

#2 of 17

Terminal tasks, earlier set

OSWorld 2.0

#7 of 10

Long desktop and web workflows

Compare Gemini 3.8 Flash side by side

Pre-filled with the two closest models. Swap any of them.

Open comparison

Sources

Building on Gemini 3.8 Flash?

Get the architecture right before the bill arrives.

We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.