Gemini 3.1 Flash-Lite
Google · Released 7 May 2026 · gemini-3.1-flash-lite
$0.25
Input price / 1M tokens
#8 of 61, cheapest first
$1.50
Output price / 1M tokens
#11 of 61, cheapest first
1.05M
Context window
#11 of 71
280
Output tokens / second
#4 of 61
16%
Humanity’s Last Exam (no tools)
#18 of 23
86.9%
GPQA Diamond
#18 of 27
76.8%
MMMU-Pro
#5 of 12
72%
LiveCodeBench
#4 of 7
Summary
Gemini 3.1 Flash-Lite is a proprietary model from Google, released on 7 May 2026.
At $0.25 input and $1.50 output per million tokens, it is #8 of 61 on input price, cheapest first.
Its 1.05M-token context window ranks #11 of 71.
Measured output speed is 280 tokens per second, #4 of 61.
Low-latency, cost-efficient multimodal model for high-frequency lightweight tasks, simple data extraction and latency-sensitive apps.
Benchmarks · 8 reported
| Benchmark | Score | Bar | Rank | Source |
|---|---|---|---|---|
| Humanity’s Last Exam (no tools)no tools, full set text + multimodal, high thinking | 16% | #18 of 23 | Google Gemini 3.1 Flash-Lite evaluation report (preview model)Self-reported | |
| GPQA Diamondno tools, high thinking | 86.9% | #18 of 27 | Google Gemini 3.1 Flash-Lite evaluation report (preview model)Self-reported | |
| MMMU-Prono tools | 76.8% | #5 of 12 | Google Gemini 3.1 Flash-Lite evaluation report (preview model)Self-reported | |
| Video-MMMU | 84.8% | – | Google Gemini 3.1 Flash-Lite evaluation report (preview model)Self-reported | |
| SimpleQA Verified | 43.3% | – | Google Gemini 3.1 Flash-Lite evaluation report (preview model)Self-reported | |
| MMMLU | 88.9% | – | Google Gemini 3.1 Flash-Lite evaluation report (preview model)Self-reported | |
| LiveCodeBenchcode generation, 175 UI problems dated 1 Jan 2025 to 1 May 2025 | 72% | #4 of 7 | Google Gemini 3.1 Flash-Lite evaluation report (preview model)Self-reported | |
| MRCR v2 (8-needle)128k average (12.3% at 1M pointwise) | 60.1% | #9 of 13 | Google Gemini 3.1 Flash-Lite evaluation report (preview model)Self-reported |
Self-reported means the lab ran the test itself. A rank appears only where several labs report the same version of a benchmark.
Similar models
Models with the most benchmarks in common with Gemini 3.1 Flash-Lite, and the closest scores.
Humanity’s Last Exam (no tools)
#18 of 23Expert questions, no search or code
GPQA Diamond
#18 of 27Graduate-level science questions
MMMU-Pro
#5 of 12Hard questions that need an image
LiveCodeBench
#4 of 7Fresh competitive programming problems
Compare Gemini 3.1 Flash-Lite side by side
Pre-filled with the two closest models. Swap any of them.
Sources
- price, cache, batch, free tier: ai.google.dev/gemini-api/docs/pricing, checked 2 Oct 2026
- context, max output, modalities, api id: ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite, checked 2 Oct 2026
- release date, GA status: ai.google.dev/gemini-api/docs/changelog, checked 2 Oct 2026
- shutdown date and replacement: ai.google.dev/gemini-api/docs/deprecations, checked 2 Oct 2026
- benchmarks: storage.googleapis.com/deepmind-media/gemini/gemini_3-1_flash-lite_model_evaluat, checked 2 Oct 2026
- output speed: artificialanalysis.ai/models/gemini-3-1-flash-lite-preview, checked 2 Oct 2026
Building on Gemini 3.1 Flash-Lite?
Get the architecture right before the bill arrives.
We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.