Skip to content
Free tool · Updated 2 Oct 2026

o3 vs gpt-oss-120b vs gpt-oss-20b

Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.

o3gpt-oss-120bgpt-oss-20b

Cheapest blended

o3

$3.50 / 1M tokens

Largest context

o3

200K tokens

Fastest output

gpt-oss-20b

180 tokens/s

Top on SWE-bench Verified

o3

69.1%

Specifications

Speco3OpenAIgpt-oss-120bOpenAIgpt-oss-20bOpenAI
ProviderOpenAIOpenAIOpenAI
Released16 Apr 20255 Aug 20255 Aug 2025
LicenceProprietaryRetires 11 Dec 2026Open weightOpen weight
Context window200K tokens131K tokens131K tokens
Max output100K tokens131K tokens131K tokens
Input / 1M$2.00*–*–*
Output / 1M$8.00––
Cached input / 1M$0.50––
Speed102 tokens/s164 tokens/s180 tokens/s
Input typestext, imagetexttext
Best forReasoningVisionReasoningBudgetBudget

* o3: Batch: $1 / $4. Flex: $1 / $4 (cached $0.25). Fast mode: $3.50 / $14. Snapshot o3-2025-04-16 shuts down 11 Dec 2026; the recommended replacement is gpt-5.6-sol. o3-pro ($20 / $80) shuts down on the same date.

* gpt-oss-120b: Open weights on Hugging Face (openai/gpt-oss-120b). Not on OpenAI's API pricing page; the price depends on the host. Runs on a single 80 GB GPU (MXFP4).

* gpt-oss-20b: Open weights on Hugging Face (openai/gpt-oss-20b). Not on OpenAI's API pricing page; the price depends on the host. Runs on devices with 16 GB of memory.

Benchmark scores

Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.

o3gpt-oss-120bgpt-oss-20b
SWE-bench Verified69.1%62.4%60.7%
AIME 202598.4%97.9%98.7%
GPQA Diamond no published score80.1%71.5%
Humanity’s Last Exam (no tools) no published score14.9%10.9%

Price per 1M tokens

o3gpt-oss-120bgpt-oss-20b
Input$2.00 no published score no published score
Output$8.00 no published score no published score
Cached input$0.50 no published score no published score

Context window and speed

o3gpt-oss-120bgpt-oss-20b
Context200K131K131K
Tokens/s102164180

Notes

o3
An o-series reasoning model for math, science, coding and visual reasoning. The docs say GPT-5 succeeded it.
gpt-oss-120b
OpenAI's larger open-weight reasoning model (117B total, 5.1B active parameters, MoE). OpenAI says it is near o4-mini on core reasoning benchmarks.
gpt-oss-20b
OpenAI's smaller open-weight reasoning model (21B total, 3.6B active parameters), for low latency and local or on-device use. OpenAI says it is similar to o3-mini.

Still torn between two?

Test them on your own prompts.

Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.