o4-mini vs gpt-oss-20b vs o3 vs gpt-oss-120b
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
o4-mini
$1.93 / 1M tokens
Largest context
o4-minio3
Tie at 200K tokens
Fastest output
gpt-oss-20b
180 tokens/s
Top on AIME 2025
o4-mini
99.5%
Specifications
| Spec | o4-miniOpenAI | gpt-oss-20bOpenAI | o3OpenAI | gpt-oss-120bOpenAI |
|---|---|---|---|---|
| Provider | OpenAI | OpenAI | OpenAI | OpenAI |
| Released | 16 Apr 2025 | 5 Aug 2025 | 16 Apr 2025 | 5 Aug 2025 |
| Licence | ProprietaryRetires 23 Oct 2026 | Open weight | ProprietaryRetires 11 Dec 2026 | Open weight |
| Context window | 200K tokens | 131K tokens | 200K tokens | 131K tokens |
| Max output | 100K tokens | 131K tokens | 100K tokens | 131K tokens |
| Input / 1M | $1.10* | –* | $2.00* | –* |
| Output / 1M | $4.40 | – | $8.00 | – |
| Cached input / 1M | $0.275 | – | $0.50 | – |
| Speed | 142 tokens/s | 180 tokens/s | 102 tokens/s | 164 tokens/s |
| Input types | text, image | text | text, image | text |
| Best for | ReasoningBudget | Budget | ReasoningVision | ReasoningBudget |
* o4-mini: Shuts down 23 Oct 2026 (both o4-mini and o4-mini-2025-04-16); the recommended replacement is gpt-5.6-terra. Batch: $0.55 / $2.20. Flex: $0.55 / $2.20 (cached $0.138). Fast mode: $2 / $8.
* gpt-oss-20b: Open weights on Hugging Face (openai/gpt-oss-20b). Not on OpenAI's API pricing page; the price depends on the host. Runs on devices with 16 GB of memory.
* o3: Batch: $1 / $4. Flex: $1 / $4 (cached $0.25). Fast mode: $3.50 / $14. Snapshot o3-2025-04-16 shuts down 11 Dec 2026; the recommended replacement is gpt-5.6-sol. o3-pro ($20 / $80) shuts down on the same date.
* gpt-oss-120b: Open weights on Hugging Face (openai/gpt-oss-120b). Not on OpenAI's API pricing page; the price depends on the host. Runs on a single 80 GB GPU (MXFP4).
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- o4-mini
- A fast, cost-efficient o-series reasoning model for coding and visual tasks. The docs say GPT-5 mini succeeded it.
- gpt-oss-20b
- OpenAI's smaller open-weight reasoning model (21B total, 3.6B active parameters), for low latency and local or on-device use. OpenAI says it is similar to o3-mini.
- o3
- An o-series reasoning model for math, science, coding and visual reasoning. The docs say GPT-5 succeeded it.
- gpt-oss-120b
- OpenAI's larger open-weight reasoning model (117B total, 5.1B active parameters, MoE). OpenAI says it is near o4-mini on core reasoning benchmarks.
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.