Skip to content

gpt-oss-20b

OpenAI · Released 5 Aug 2025 · gpt-oss-20b

Compare this model
Open weightBudget

–

Input price / 1M tokens

–

Output price / 1M tokens

131K

Context window

#68 of 71

180

Output tokens / second

#15 of 61

60.7%

SWE-bench Verified

#10 of 12

71.5%

GPQA Diamond

#24 of 27

98.7%

AIME 2025

#2 of 7

10.9%

Humanity’s Last Exam (no tools)

#21 of 23

Summary

gpt-oss-20b is an open-weight model from OpenAI, released on 5 Aug 2025.

Its 131K-token context window ranks #68 of 71.

Measured output speed is 180 tokens per second, #15 of 61.

OpenAI's smaller open-weight reasoning model (21B total, 3.6B active parameters), for low latency and local or on-device use. OpenAI says it is similar to o3-mini.

Benchmarks · 5 reported

BenchmarkScoreBarRankSource
SWE-bench Verifiedhigh reasoning effort60.7%#10 of 12OpenAI gpt-oss model cardSelf-reported
GPQA Diamondno tools, high reasoning effort71.5%#24 of 27OpenAI gpt-oss model cardSelf-reported
AIME 2025with tools, high reasoning effort98.7%#2 of 7OpenAI gpt-oss model cardSelf-reported
Humanity’s Last Exam (no tools)no tools, high reasoning effort10.9%#21 of 23OpenAI gpt-oss model cardSelf-reported
MMLUhigh reasoning effort85.3%–OpenAI gpt-oss model cardSelf-reported

Self-reported means the lab ran the test itself. A rank appears only where several labs report the same version of a benchmark.

Similar models

Models with the most benchmarks in common with gpt-oss-20b, and the closest scores.

Compare 4 side by side

SWE-bench Verified

#10 of 12

Real GitHub issues, fixed and tested

GPQA Diamond

#24 of 27

Graduate-level science questions

AIME 2025

#2 of 7

Competition mathematics

Humanity’s Last Exam (no tools)

#21 of 23

Expert questions, no search or code

Compare gpt-oss-20b side by side

Pre-filled with the two closest models. Swap any of them.

Open comparison

Sources

Building on gpt-oss-20b?

Get the architecture right before the bill arrives.

We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.