Skip to content

Gemini 3.1 Pro

Google · Released 19 Feb 2026 · gemini-3.1-pro-preview

Compare this model
ProprietaryPreviewReasoningCodingAgentsMultimodal

$2.00

Input price / 1M tokens

#36 of 61, cheapest first

$12.00

Output price / 1M tokens

#44 of 61, cheapest first

1.05M

Context window

#11 of 71

114

Output tokens / second

#26 of 61

44.4%

Humanity’s Last Exam (no tools)

#4 of 23

77.1%

ARC-AGI-2

#3 of 6

94.3%

GPQA Diamond

#3 of 27

68.5%

Terminal-Bench 2.0

#3 of 7

Summary

Gemini 3.1 Pro is a proprietary model from Google, released on 19 Feb 2026.

At $2.00 input and $12.00 output per million tokens, it is #36 of 61 on input price, cheapest first.

Its 1.05M-token context window ranks #11 of 71.

Measured output speed is 114 tokens per second, #26 of 61.

Google's current Pro model, still preview. Built for multimodal understanding, agentic work and vibe-coding, with a focus on software engineering and multi-step tool use.

Benchmarks · 15 reported

BenchmarkScoreBarRankSource
Humanity’s Last Exam (no tools)no tools, full set text + multimodal, thinking high44.4%#4 of 23Google Gemini 3.1 Pro evaluation reportSelf-reported
Humanity's Last Examsearch (blocklist) + code, thinking high51.4%–Google Gemini 3.1 Pro evaluation reportSelf-reported
ARC-AGI-2ARC Prize Verified, semi-private set77.1%#3 of 6ARC Prize, cited in Google's evaluation report
GPQA Diamondno tools94.3%#3 of 27Google Gemini 3.1 Pro evaluation reportSelf-reported
Terminal-Bench 2.0Terminus 2 harness68.5%#3 of 7Google Gemini 3.1 Pro evaluation reportSelf-reported
SWE-bench Verifiedsingle attempt, average of 10 runs; includes a 0.6-point adjustment for 3 broken harness items80.6%#1 of 12Google Gemini 3.1 Pro evaluation reportSelf-reported
SWE-bench Prosingle attempt, average of 5 runs54.2%#14 of 16Google Gemini 3.1 Pro evaluation reportSelf-reported
LiveCodeBench Propublic leaderboard2,887Elo-style–LiveCodeBench Pro leaderboard, cited in Google's evaluation report
APEX-Agents33.5%–Google Gemini 3.1 Pro evaluation reportSelf-reported
τ²-benchretail (99.3% telecom)90.8%#2 of 6Google Gemini 3.1 Pro evaluation reportSelf-reported
MCP Atlaspublic set, sourced from Turing69.2%–Turing, cited in Google's evaluation report
BrowseCompDeep Research with search, Python and browsing85.9%#3 of 7Google Gemini 3.1 Pro evaluation reportSelf-reported
MMMU-Prono tools80.5%#3 of 12Google Gemini 3.1 Pro evaluation reportSelf-reported
MMMLU92.6%–Google Gemini 3.1 Pro evaluation reportSelf-reported
MRCR v2 (8-needle)128k average (26.3% at 1M pointwise)84.9%#3 of 13Google Gemini 3.1 Pro evaluation reportSelf-reported

Self-reported means the lab ran the test itself. A rank appears only where several labs report the same version of a benchmark.

Similar models

Models with the most benchmarks in common with Gemini 3.1 Pro, and the closest scores.

Compare 4 side by side

Humanity’s Last Exam (no tools)

#4 of 23

Expert questions, no search or code

ARC-AGI-2

#3 of 6

Learning a new rule from examples

GPQA Diamond

#3 of 27

Graduate-level science questions

Terminal-Bench 2.0

#3 of 7

Terminal tasks, 2025 set

Compare Gemini 3.1 Pro side by side

Pre-filled with the two closest models. Swap any of them.

Open comparison

Sources

Building on Gemini 3.1 Pro?

Get the architecture right before the bill arrives.

We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.