GPT-5 vs Gemini 3 Flash vs Gemini 3.1 Pro
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
Gemini 3 Flash
$1.13 / 1M tokens
Largest context
Gemini 3 FlashGemini 3.1 Pro
Tie at 1.05M tokens
Fastest output
Gemini 3 Flash
182 tokens/s
Top on SWE-bench Verified
Gemini 3.1 Pro
80.6%
Specifications
| Spec | GPT-5OpenAI | Gemini 3 FlashGoogle | Gemini 3.1 ProGoogle |
|---|---|---|---|
| Provider | OpenAI | ||
| Released | 7 Aug 2025 | 17 Dec 2025 | 19 Feb 2026 |
| Licence | ProprietaryRetires 11 Dec 2026 | ProprietaryPreview | ProprietaryPreview |
| Context window | 400K tokens | 1.05M tokens | 1.05M tokens |
| Max output | 128K tokens | 66K tokens | 66K tokens |
| Input / 1M | $1.25* | $0.50* | $2.00* |
| Output / 1M | $10.00 | $3.00 | $12.00 |
| Cached input / 1M | $0.125 | $0.05 | $0.20 |
| Speed | 75 tokens/s | 182 tokens/s | 114 tokens/s |
| Input types | text, image | text, image, video, audio, pdf | text, image, video, audio, pdf |
| Best for | Coding | Budget | ReasoningCodingAgentsMultimodal |
* GPT-5: There is no separate cache-write charge. Batch and Flex: $0.625 / $5. Fast mode: $2.50 / $20. Snapshot gpt-5-2025-08-07 shuts down 11 Dec 2026; the recommended replacement is gpt-5.6-sol.
* Gemini 3 Flash: Audio input costs $1.00, or $0.10 cached. Cache storage $1.00 per 1M tokens per hour. Google calls it a legacy model; the deprecations page names gemini-3.6-flash as the replacement, with no shutdown date yet.
* Gemini 3.1 Pro: Still preview. A separate endpoint, gemini-3.1-pro-preview-customtools, has the same price. Cache storage $4.50 per 1M tokens per hour. Priority is $3.60 / $21.60 up to 200K.
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- GPT-5
- The previous-generation reasoning model for coding and agentic tasks, launched August 2025. OpenAI now points users to newer models.
- Gemini 3 Flash
- Legacy Gemini 3 preview Flash model giving baseline speed and intelligence. It is still served, but newer GA Flash models replace it.
- Gemini 3.1 Pro
- Google's current Pro model, still preview. Built for multimodal understanding, agentic work and vibe-coding, with a focus on software engineering and multi-step tool use.
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.