Gemini 3.1 Pro vs Claude Opus 5.5
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
Gemini 3.1 Pro
$4.50 / 1M tokens
Largest context
Gemini 3.1 Pro
1.05M tokens
Fastest output
Gemini 3.1 Pro
114 tokens/s
Top benchmark
–
No benchmark they share
Specifications
| Spec | Gemini 3.1 ProGoogle | Claude Opus 5.5Anthropic |
|---|---|---|
| Provider | Anthropic | |
| Released | 19 Feb 2026 | 22 Sep 2026 |
| Licence | ProprietaryPreview | Proprietary |
| Context window | 1.05M tokens | 1M tokens |
| Max output | 66K tokens | 128K tokens |
| Input / 1M | $2.00* | $4.00* |
| Output / 1M | $12.00 | $20.00 |
| Cached input / 1M | $0.20 | $0.20 |
| Speed | 114 tokens/s | 93 tokens/s |
| Input types | text, image, video, audio, pdf | text, image |
| Best for | ReasoningCodingAgentsMultimodal | CodingAgentsReasoningLong context |
* Gemini 3.1 Pro: Still preview. A separate endpoint, gemini-3.1-pro-preview-customtools, has the same price. Cache storage $4.50 per 1M tokens per hour. Priority is $3.60 / $21.60 up to 200K.
* Claude Opus 5.5: 20% cheaper per token than Opus 5 ($5/$25). Cache reads are 0.05x base input ($0.20). Fast mode (research preview, Claude API only): $8 input / $40 output. Uses the newer tokenizer (about 30% more tokens than pre-Opus 4.7 models). Safety classifiers can refuse; since 24 Sep 2026 some pre-output refusal categories are billed. Minimum cacheable prompt: 512 tokens.
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- Gemini 3.1 Pro
- Google's current Pro model, still preview. Built for multimodal understanding, agentic work and vibe-coding, with a focus on software engineering and multi-step tool use.
- Claude Opus 5.5
- Anthropic's recommended starting model for most workloads, built for long-running agentic coding and knowledge work. Anthropic says it performs at the level of Claude Fable 5.1 on most work at lower cost.
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.