Gemini 2.5 Flash
Google · Released 17 Jun 2025 · gemini-2.5-flash
$0.30
Input price / 1M tokens
$2.50
Output price / 1M tokens
1.05M
Context window
#11 of 71
191
Output tokens / second
#11 of 61
11%
Humanity’s Last Exam (no tools)
#20 of 23
82.8%
GPQA Diamond
#21 of 27
72%
AIME 2025
#7 of 7
66.7%
MMMU-Pro
#8 of 12
Summary
Gemini 2.5 Flash is a proprietary model from Google, released on 17 Jun 2025.
Its 1.05M-token context window ranks #11 of 71.
Measured output speed is 191 tokens per second, #11 of 61.
Gemini 2.5 hybrid reasoning model with thinking budgets and a 1M context, aimed at large-scale, low-latency work. Closed to new users.
Benchmarks · 7 reported
| Benchmark | Score | Bar | Rank | Source |
|---|---|---|---|---|
| Humanity’s Last Exam (no tools)no tools, thinking | 11% | #20 of 23 | Google Gemini 3 Flash model card (Gemini 2.5 Flash column, Dec 2025)Self-reported | |
| GPQA Diamondno tools | 82.8% | #21 of 27 | Google Gemini 3 Flash model card (Gemini 2.5 Flash column, Dec 2025)Self-reported | |
| AIME 2025no tools (75.7% with code execution) | 72% | #7 of 7 | Google Gemini 3 Flash model card (Gemini 2.5 Flash column, Dec 2025)Self-reported | |
| MMMU-Pro | 66.7% | #8 of 12 | Google Gemini 3 Flash model card (Gemini 2.5 Flash column, Dec 2025)Self-reported | |
| SWE-bench Verifiedsingle attempt | 60.4% | #11 of 12 | Google Gemini 3 Flash model card (Gemini 2.5 Flash column, Dec 2025)Self-reported | |
| Terminal-Bench 2.0Terminus 2 harness | 16.9% | #7 of 7 | Google Gemini 3 Flash model card (Gemini 2.5 Flash column, Dec 2025)Self-reported | |
| MRCR v2 (8-needle)128k average (21.0% at 1M pointwise) | 54.3% | #11 of 13 | Google Gemini 3 Flash model card (Gemini 2.5 Flash column, Dec 2025)Self-reported |
Self-reported means the lab ran the test itself. A rank appears only where several labs report the same version of a benchmark.
Similar models
Models with the most benchmarks in common with Gemini 2.5 Flash, and the closest scores.
Humanity’s Last Exam (no tools)
#20 of 23Expert questions, no search or code
GPQA Diamond
#21 of 27Graduate-level science questions
- Gemini 2.5 Flash82.8
- Gemini 2.5 Pro86.4
- Gemini 3 Flash90.4
- Gemini 3.1 Pro94.3
- Gemma 4 26B A4B82.3
- Gemini 3.1 Flash-Lite86.9
MMMU-Pro
#8 of 12Hard questions that need an image
Compare Gemini 2.5 Flash side by side
Pre-filled with the two closest models. Swap any of them.
Sources
- price, cache, batch, free tier: ai.google.dev/gemini-api/docs/pricing, checked 2 Oct 2026
- context, max output, knowledge cutoff, modalities, access limit: ai.google.dev/gemini-api/docs/models/gemini-2.5-flash, checked 2 Oct 2026
- release date, no shutdown date: ai.google.dev/gemini-api/docs/deprecations, checked 2 Oct 2026
- benchmarks: storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Flash-Model-Card.pdf, checked 2 Oct 2026
- output speed: artificialanalysis.ai/models/gemini-2-5-flash, checked 2 Oct 2026
Building on Gemini 2.5 Flash?
Get the architecture right before the bill arrives.
We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.