Command A+ vs Gemma 4 31B vs Gemini 3.1 Pro
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
Gemini 3.1 Pro
$4.50 / 1M tokens
Largest context
Gemini 3.1 Pro
1.05M tokens
Fastest output
Command A+
211 tokens/s
Top on τ²-bench
Gemini 3.1 Pro
90.8%
Specifications
| Spec | Command A+Cohere | Gemma 4 31BGoogle | Gemini 3.1 ProGoogle |
|---|---|---|---|
| Provider | Cohere | ||
| Released | 20 May 2026 | 31 Mar 2026 | 19 Feb 2026 |
| Licence | Open weight | Open weight | ProprietaryPreview |
| Context window | 128K tokens | 262K tokens | 1.05M tokens |
| Max output | 64K tokens | – | 66K tokens |
| Input / 1M | –* | –* | $2.00* |
| Output / 1M | – | – | $12.00 |
| Cached input / 1M | – | – | $0.20 |
| Speed | 211 tokens/s | 35 tokens/s | 114 tokens/s |
| Input types | text, image | text, image | text, image, video, audio, pdf |
| Best for | AgentsReasoningMultimodal | ReasoningCodingVision | ReasoningCodingAgentsMultimodal |
* Command A+: Cohere publishes no per-token rate. Its pricing page lists Command A+ as "Free" ($0 API key, $0 model download); production rate limits are "contact sales", and dedicated hosting is via Model Vault. Third-party sites quote other prices that Cohere does not show.
* Gemma 4 31B: Open weights. On the Gemini API Gemma 4 has only a free tier; the paid tier is listed as 'Not available', so there is no per-token price.
* Gemini 3.1 Pro: Still preview. A separate endpoint, gemini-3.1-pro-preview-customtools, has the same price. Cache storage $4.50 per 1M tokens per hour. Priority is $3.60 / $21.60 up to 200K.
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- Command A+
- Cohere's first mixture-of-experts model (25B active, 218B total) and last of the Command A family, combining vision input, agentic tool use, reasoning and translation across 48 languages; it runs on 2 H100s or 1 B200.
- Gemma 4 31B
- Largest dense Gemma 4 open model (30.7B parameters), with a thinking mode, function calling and 256K context. Video is handled as frame sequences.
- Gemini 3.1 Pro
- Google's current Pro model, still preview. Built for multimodal understanding, agentic work and vibe-coding, with a focus on software engineering and multi-step tool use.
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.