Skip to content

gpt-oss-120b

OpenAI · Released 5 Aug 2025 · gpt-oss-120b

Compare this model
Open weightReasoningBudget

–

Input price / 1M tokens

–

Output price / 1M tokens

131K

Context window

#68 of 71

164

Output tokens / second

#17 of 61

62.4%

SWE-bench Verified

#9 of 12

80.1%

GPQA Diamond

#23 of 27

97.9%

AIME 2025

#4 of 7

14.9%

Humanity’s Last Exam (no tools)

#19 of 23

Summary

gpt-oss-120b is an open-weight model from OpenAI, released on 5 Aug 2025.

Its 131K-token context window ranks #68 of 71.

Measured output speed is 164 tokens per second, #17 of 61.

OpenAI's larger open-weight reasoning model (117B total, 5.1B active parameters, MoE). OpenAI says it is near o4-mini on core reasoning benchmarks.

Benchmarks · 5 reported

BenchmarkScoreBarRankSource
SWE-bench Verifiedhigh reasoning effort62.4%#9 of 12OpenAI gpt-oss model cardSelf-reported
GPQA Diamondno tools, high reasoning effort80.1%#23 of 27OpenAI gpt-oss model cardSelf-reported
AIME 2025with tools, high reasoning effort97.9%#4 of 7OpenAI gpt-oss model cardSelf-reported
Humanity’s Last Exam (no tools)no tools, high reasoning effort14.9%#19 of 23OpenAI gpt-oss model cardSelf-reported
MMLUhigh reasoning effort90%–OpenAI gpt-oss model cardSelf-reported

Self-reported means the lab ran the test itself. A rank appears only where several labs report the same version of a benchmark.

Similar models

Models with the most benchmarks in common with gpt-oss-120b, and the closest scores.

Compare 4 side by side

SWE-bench Verified

#9 of 12

Real GitHub issues, fixed and tested

GPQA Diamond

#23 of 27

Graduate-level science questions

AIME 2025

#4 of 7

Competition mathematics

Humanity’s Last Exam (no tools)

#19 of 23

Expert questions, no search or code

Compare gpt-oss-120b side by side

Pre-filled with the two closest models. Swap any of them.

Open comparison

Sources

Building on gpt-oss-120b?

Get the architecture right before the bill arrives.

We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.