Skip to content
Free tool · Updated 2 Oct 2026

gpt-oss-20b vs gpt-oss-120b vs Gemini 2.5 Pro

Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.

gpt-oss-20bgpt-oss-120bGemini 2.5 Pro

Cheapest blended

Gemini 2.5 Pro

$3.44 / 1M tokens

Largest context

Gemini 2.5 Pro

1.05M tokens

Fastest output

gpt-oss-20b

180 tokens/s

Top on SWE-bench Verified

gpt-oss-120b

62.4%

Specifications

Specgpt-oss-20bOpenAIgpt-oss-120bOpenAIGemini 2.5 ProGoogle
ProviderOpenAIOpenAIGoogle
Released5 Aug 20255 Aug 202517 Jun 2025
LicenceOpen weightOpen weightProprietaryExisting users only
Context window131K tokens131K tokens1.05M tokens
Max output131K tokens131K tokens66K tokens
Input / 1M–*–*$1.25*
Output / 1M––$10.00
Cached input / 1M––$0.125
Speed180 tokens/s164 tokens/s122 tokens/s
Input typestexttexttext, image, video, audio, pdf
Best forBudgetReasoningBudgetCodingReasoning

* gpt-oss-20b: Open weights on Hugging Face (openai/gpt-oss-20b). Not on OpenAI's API pricing page; the price depends on the host. Runs on devices with 16 GB of memory.

* gpt-oss-120b: Open weights on Hugging Face (openai/gpt-oss-120b). Not on OpenAI's API pricing page; the price depends on the host. Runs on a single 80 GB GPU (MXFP4).

* Gemini 2.5 Pro: Since 18 Sep 2026 Google has limited 2.5 access to users who actively used these models before. They are not deprecated and have no shutdown date. Cache storage $4.50 per 1M tokens per hour.

Benchmark scores

Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.

gpt-oss-20bgpt-oss-120bGemini 2.5 Pro
SWE-bench Verified60.7%62.4%59.6%
GPQA Diamond71.5%80.1%86.4%
AIME 202598.7%97.9%88%
Humanity’s Last Exam (no tools)10.9%14.9%21.6%

Price per 1M tokens

gpt-oss-20bgpt-oss-120bGemini 2.5 Pro
Input no published score no published score$1.25
Output no published score no published score$10.00
Cached input no published score no published score$0.125

Context window and speed

gpt-oss-20bgpt-oss-120bGemini 2.5 Pro
Context131K131K1.05M
Tokens/s180164122

Notes

gpt-oss-20b
OpenAI's smaller open-weight reasoning model (21B total, 3.6B active parameters), for low latency and local or on-device use. OpenAI says it is similar to o3-mini.
gpt-oss-120b
OpenAI's larger open-weight reasoning model (117B total, 5.1B active parameters, MoE). OpenAI says it is near o4-mini on core reasoning benchmarks.
Gemini 2.5 Pro
Gemini 2.5 thinking model for complex reasoning in code, math and STEM. It is still served but closed to new users; Google points new projects to 3.5 Flash-Lite or 3.8 Flash.

Still torn between two?

Test them on your own prompts.

Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.