o3 vs gpt-oss-120b vs gpt-oss-20b vs Gemini 3 Flash
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
Gemini 3 Flash
$1.13 / 1M tokens
Largest context
Gemini 3 Flash
1.05M tokens
Fastest output
Gemini 3 Flash
182 tokens/s
Top on SWE-bench Verified
Gemini 3 Flash
78%
Specifications
| Spec | o3OpenAI | gpt-oss-120bOpenAI | gpt-oss-20bOpenAI | Gemini 3 FlashGoogle |
|---|---|---|---|---|
| Provider | OpenAI | OpenAI | OpenAI | |
| Released | 16 Apr 2025 | 5 Aug 2025 | 5 Aug 2025 | 17 Dec 2025 |
| Licence | ProprietaryRetires 11 Dec 2026 | Open weight | Open weight | ProprietaryPreview |
| Context window | 200K tokens | 131K tokens | 131K tokens | 1.05M tokens |
| Max output | 100K tokens | 131K tokens | 131K tokens | 66K tokens |
| Input / 1M | $2.00* | –* | –* | $0.50* |
| Output / 1M | $8.00 | – | – | $3.00 |
| Cached input / 1M | $0.50 | – | – | $0.05 |
| Speed | 102 tokens/s | 164 tokens/s | 180 tokens/s | 182 tokens/s |
| Input types | text, image | text | text | text, image, video, audio, pdf |
| Best for | ReasoningVision | ReasoningBudget | Budget | Budget |
* o3: Batch: $1 / $4. Flex: $1 / $4 (cached $0.25). Fast mode: $3.50 / $14. Snapshot o3-2025-04-16 shuts down 11 Dec 2026; the recommended replacement is gpt-5.6-sol. o3-pro ($20 / $80) shuts down on the same date.
* gpt-oss-120b: Open weights on Hugging Face (openai/gpt-oss-120b). Not on OpenAI's API pricing page; the price depends on the host. Runs on a single 80 GB GPU (MXFP4).
* gpt-oss-20b: Open weights on Hugging Face (openai/gpt-oss-20b). Not on OpenAI's API pricing page; the price depends on the host. Runs on devices with 16 GB of memory.
* Gemini 3 Flash: Audio input costs $1.00, or $0.10 cached. Cache storage $1.00 per 1M tokens per hour. Google calls it a legacy model; the deprecations page names gemini-3.6-flash as the replacement, with no shutdown date yet.
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- o3
- An o-series reasoning model for math, science, coding and visual reasoning. The docs say GPT-5 succeeded it.
- gpt-oss-120b
- OpenAI's larger open-weight reasoning model (117B total, 5.1B active parameters, MoE). OpenAI says it is near o4-mini on core reasoning benchmarks.
- gpt-oss-20b
- OpenAI's smaller open-weight reasoning model (21B total, 3.6B active parameters), for low latency and local or on-device use. OpenAI says it is similar to o3-mini.
- Gemini 3 Flash
- Legacy Gemini 3 preview Flash model giving baseline speed and intelligence. It is still served, but newer GA Flash models replace it.
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.