Skip to content
Free tool · Updated 2 Oct 2026

o3 vs gpt-oss-120b vs gpt-oss-20b vs Gemini 3 Flash

Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.

o3gpt-oss-120bgpt-oss-20bGemini 3 Flash

Cheapest blended

Gemini 3 Flash

$1.13 / 1M tokens

Largest context

Gemini 3 Flash

1.05M tokens

Fastest output

Gemini 3 Flash

182 tokens/s

Top on SWE-bench Verified

Gemini 3 Flash

78%

Specifications

Speco3OpenAIgpt-oss-120bOpenAIgpt-oss-20bOpenAIGemini 3 FlashGoogle
ProviderOpenAIOpenAIOpenAIGoogle
Released16 Apr 20255 Aug 20255 Aug 202517 Dec 2025
LicenceProprietaryRetires 11 Dec 2026Open weightOpen weightProprietaryPreview
Context window200K tokens131K tokens131K tokens1.05M tokens
Max output100K tokens131K tokens131K tokens66K tokens
Input / 1M$2.00*–*–*$0.50*
Output / 1M$8.00––$3.00
Cached input / 1M$0.50––$0.05
Speed102 tokens/s164 tokens/s180 tokens/s182 tokens/s
Input typestext, imagetexttexttext, image, video, audio, pdf
Best forReasoningVisionReasoningBudgetBudgetBudget

* o3: Batch: $1 / $4. Flex: $1 / $4 (cached $0.25). Fast mode: $3.50 / $14. Snapshot o3-2025-04-16 shuts down 11 Dec 2026; the recommended replacement is gpt-5.6-sol. o3-pro ($20 / $80) shuts down on the same date.

* gpt-oss-120b: Open weights on Hugging Face (openai/gpt-oss-120b). Not on OpenAI's API pricing page; the price depends on the host. Runs on a single 80 GB GPU (MXFP4).

* gpt-oss-20b: Open weights on Hugging Face (openai/gpt-oss-20b). Not on OpenAI's API pricing page; the price depends on the host. Runs on devices with 16 GB of memory.

* Gemini 3 Flash: Audio input costs $1.00, or $0.10 cached. Cache storage $1.00 per 1M tokens per hour. Google calls it a legacy model; the deprecations page names gemini-3.6-flash as the replacement, with no shutdown date yet.

Benchmark scores

Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.

o3gpt-oss-120bgpt-oss-20bGemini 3 Flash
SWE-bench Verified69.1%62.4%60.7%78%
AIME 202598.4%97.9%98.7%95.2%
GPQA Diamond no published score80.1%71.5%90.4%
Humanity’s Last Exam (no tools) no published score14.9%10.9%33.7%

Price per 1M tokens

o3gpt-oss-120bgpt-oss-20bGemini 3 Flash
Input$2.00 no published score no published score$0.50
Output$8.00 no published score no published score$3.00
Cached input$0.50 no published score no published score$0.05

Context window and speed

o3gpt-oss-120bgpt-oss-20bGemini 3 Flash
Context200K131K131K1.05M
Tokens/s102164180182

Notes

o3
An o-series reasoning model for math, science, coding and visual reasoning. The docs say GPT-5 succeeded it.
gpt-oss-120b
OpenAI's larger open-weight reasoning model (117B total, 5.1B active parameters, MoE). OpenAI says it is near o4-mini on core reasoning benchmarks.
gpt-oss-20b
OpenAI's smaller open-weight reasoning model (21B total, 3.6B active parameters), for low latency and local or on-device use. OpenAI says it is similar to o3-mini.
Gemini 3 Flash
Legacy Gemini 3 preview Flash model giving baseline speed and intelligence. It is still served, but newer GA Flash models replace it.

Still torn between two?

Test them on your own prompts.

Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.