gpt-oss-20b vs gpt-oss-120b vs Gemini 2.5 Pro
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
Gemini 2.5 Pro
$3.44 / 1M tokens
Largest context
Gemini 2.5 Pro
1.05M tokens
Fastest output
gpt-oss-20b
180 tokens/s
Top on SWE-bench Verified
gpt-oss-120b
62.4%
Specifications
| Spec | gpt-oss-20bOpenAI | gpt-oss-120bOpenAI | Gemini 2.5 ProGoogle |
|---|---|---|---|
| Provider | OpenAI | OpenAI | |
| Released | 5 Aug 2025 | 5 Aug 2025 | 17 Jun 2025 |
| Licence | Open weight | Open weight | ProprietaryExisting users only |
| Context window | 131K tokens | 131K tokens | 1.05M tokens |
| Max output | 131K tokens | 131K tokens | 66K tokens |
| Input / 1M | –* | –* | $1.25* |
| Output / 1M | – | – | $10.00 |
| Cached input / 1M | – | – | $0.125 |
| Speed | 180 tokens/s | 164 tokens/s | 122 tokens/s |
| Input types | text | text | text, image, video, audio, pdf |
| Best for | Budget | ReasoningBudget | CodingReasoning |
* gpt-oss-20b: Open weights on Hugging Face (openai/gpt-oss-20b). Not on OpenAI's API pricing page; the price depends on the host. Runs on devices with 16 GB of memory.
* gpt-oss-120b: Open weights on Hugging Face (openai/gpt-oss-120b). Not on OpenAI's API pricing page; the price depends on the host. Runs on a single 80 GB GPU (MXFP4).
* Gemini 2.5 Pro: Since 18 Sep 2026 Google has limited 2.5 access to users who actively used these models before. They are not deprecated and have no shutdown date. Cache storage $4.50 per 1M tokens per hour.
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- gpt-oss-20b
- OpenAI's smaller open-weight reasoning model (21B total, 3.6B active parameters), for low latency and local or on-device use. OpenAI says it is similar to o3-mini.
- gpt-oss-120b
- OpenAI's larger open-weight reasoning model (117B total, 5.1B active parameters, MoE). OpenAI says it is near o4-mini on core reasoning benchmarks.
- Gemini 2.5 Pro
- Gemini 2.5 thinking model for complex reasoning in code, math and STEM. It is still served but closed to new users; Google points new projects to 3.5 Flash-Lite or 3.8 Flash.
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.