Skip to content
Free tool · Updated 2 Oct 2026

o4-mini vs gpt-oss-20b vs o3

Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.

o4-minigpt-oss-20bo3

Cheapest blended

o4-mini

$1.93 / 1M tokens

Largest context

o4-minio3

Tie at 200K tokens

Fastest output

gpt-oss-20b

180 tokens/s

Top on AIME 2025

o4-mini

99.5%

Specifications

Speco4-miniOpenAIgpt-oss-20bOpenAIo3OpenAI
ProviderOpenAIOpenAIOpenAI
Released16 Apr 20255 Aug 202516 Apr 2025
LicenceProprietaryRetires 23 Oct 2026Open weightProprietaryRetires 11 Dec 2026
Context window200K tokens131K tokens200K tokens
Max output100K tokens131K tokens100K tokens
Input / 1M$1.10*–*$2.00*
Output / 1M$4.40–$8.00
Cached input / 1M$0.275–$0.50
Speed142 tokens/s180 tokens/s102 tokens/s
Input typestext, imagetexttext, image
Best forReasoningBudgetBudgetReasoningVision

* o4-mini: Shuts down 23 Oct 2026 (both o4-mini and o4-mini-2025-04-16); the recommended replacement is gpt-5.6-terra. Batch: $0.55 / $2.20. Flex: $0.55 / $2.20 (cached $0.138). Fast mode: $2 / $8.

* gpt-oss-20b: Open weights on Hugging Face (openai/gpt-oss-20b). Not on OpenAI's API pricing page; the price depends on the host. Runs on devices with 16 GB of memory.

* o3: Batch: $1 / $4. Flex: $1 / $4 (cached $0.25). Fast mode: $3.50 / $14. Snapshot o3-2025-04-16 shuts down 11 Dec 2026; the recommended replacement is gpt-5.6-sol. o3-pro ($20 / $80) shuts down on the same date.

Benchmark scores

Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.

o4-minigpt-oss-20bo3
AIME 202599.5%98.7%98.4%
SWE-bench Verified no published score60.7%69.1%
GPQA Diamond no published score71.5% no published score
Humanity’s Last Exam (no tools) no published score10.9% no published score

Price per 1M tokens

o4-minigpt-oss-20bo3
Input$1.10 no published score$2.00
Output$4.40 no published score$8.00
Cached input$0.275 no published score$0.50

Context window and speed

o4-minigpt-oss-20bo3
Context200K131K200K
Tokens/s142180102

Notes

o4-mini
A fast, cost-efficient o-series reasoning model for coding and visual tasks. The docs say GPT-5 mini succeeded it.
gpt-oss-20b
OpenAI's smaller open-weight reasoning model (21B total, 3.6B active parameters), for low latency and local or on-device use. OpenAI says it is similar to o3-mini.
o3
An o-series reasoning model for math, science, coding and visual reasoning. The docs say GPT-5 succeeded it.

Still torn between two?

Test them on your own prompts.

Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.