Gemini 3.8 Flash vs Claude Haiku 4.5
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
Gemini 3.8 Flash
$1.50 / 1M tokens
Largest context
Gemini 3.8 Flash
1.05M tokens
Fastest output
Gemini 3.8 Flash
249 tokens/s
Top benchmark
–
No benchmark they share
Specifications
| Spec | Gemini 3.8 FlashGoogle | Claude Haiku 4.5Anthropic |
|---|---|---|
| Provider | Anthropic | |
| Released | 2 Sep 2026 | 15 Oct 2025 |
| Licence | Proprietary | Proprietary |
| Context window | 1.05M tokens | 200K tokens |
| Max output | 66K tokens | 64K tokens |
| Input / 1M | $0.75* | $1.00* |
| Output / 1M | $3.75 | $5.00 |
| Cached input / 1M | $0.075 | $0.10 |
| Speed | 249 tokens/s | 91 tokens/s |
| Input types | text, image, video, audio, pdf | text, image |
| Best for | CodingAgentsReasoningMultimodal | BudgetCoding |
* Gemini 3.8 Flash: Introductory price through 31 Dec 2026. From 1 Jan 2027: $1.50 input, $7.50 output, $0.15 cached input. Cache storage $0.50 per 1M tokens per hour, rising to $1.00 in 2027. Google bills no separate cache write. One flat price up to the full 1M context. Flex is also 50% off; Priority is $1.35 / $6.75 ($2.70 / $13.50 from 2027).
* Claude Haiku 4.5: Cheapest current Claude model. Uses the previous tokenizer. Minimum cacheable prompt: 4,096 tokens. Retirement commitment is only 'not sooner than 15 Oct 2026'; no deprecation notice yet.
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- Gemini 3.8 Flash
- Google's most intelligent Flash model, built for long-horizon software engineering, autonomous agents and complex enterprise workflows at Flash speed and cost.
- Claude Haiku 4.5
- Anthropic's fastest model with near-frontier intelligence; Anthropic positions it as giving coding performance similar to Claude Sonnet 4 at one-third the cost and more than twice the speed.
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.