o3 vs gpt-oss-120b vs gpt-oss-20b
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
o3
$3.50 / 1M tokens
Largest context
o3
200K tokens
Fastest output
gpt-oss-20b
180 tokens/s
Top on SWE-bench Verified
o3
69.1%
Specifications
| Spec | o3OpenAI | gpt-oss-120bOpenAI | gpt-oss-20bOpenAI |
|---|---|---|---|
| Provider | OpenAI | OpenAI | OpenAI |
| Released | 16 Apr 2025 | 5 Aug 2025 | 5 Aug 2025 |
| Licence | ProprietaryRetires 11 Dec 2026 | Open weight | Open weight |
| Context window | 200K tokens | 131K tokens | 131K tokens |
| Max output | 100K tokens | 131K tokens | 131K tokens |
| Input / 1M | $2.00* | –* | –* |
| Output / 1M | $8.00 | – | – |
| Cached input / 1M | $0.50 | – | – |
| Speed | 102 tokens/s | 164 tokens/s | 180 tokens/s |
| Input types | text, image | text | text |
| Best for | ReasoningVision | ReasoningBudget | Budget |
* o3: Batch: $1 / $4. Flex: $1 / $4 (cached $0.25). Fast mode: $3.50 / $14. Snapshot o3-2025-04-16 shuts down 11 Dec 2026; the recommended replacement is gpt-5.6-sol. o3-pro ($20 / $80) shuts down on the same date.
* gpt-oss-120b: Open weights on Hugging Face (openai/gpt-oss-120b). Not on OpenAI's API pricing page; the price depends on the host. Runs on a single 80 GB GPU (MXFP4).
* gpt-oss-20b: Open weights on Hugging Face (openai/gpt-oss-20b). Not on OpenAI's API pricing page; the price depends on the host. Runs on devices with 16 GB of memory.
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- o3
- An o-series reasoning model for math, science, coding and visual reasoning. The docs say GPT-5 succeeded it.
- gpt-oss-120b
- OpenAI's larger open-weight reasoning model (117B total, 5.1B active parameters, MoE). OpenAI says it is near o4-mini on core reasoning benchmarks.
- gpt-oss-20b
- OpenAI's smaller open-weight reasoning model (21B total, 3.6B active parameters), for low latency and local or on-device use. OpenAI says it is similar to o3-mini.
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.