Grok 4.7 vs Gemini 3.8 Flash vs GPT-6 Astra
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
Gemini 3.8 Flash
$1.50 / 1M tokens
Largest context
GPT-6 Astra
1.05M tokens
Fastest output
Gemini 3.8 Flash
249 tokens/s
Top on DeepSWE v1.1
GPT-6 Astra
74.1%
Specifications
| Spec | Grok 4.7xAI | Gemini 3.8 FlashGoogle | GPT-6 AstraOpenAI |
|---|---|---|---|
| Provider | xAI | OpenAI | |
| Released | 21 Sep 2026 | 2 Sep 2026 | 3 Sep 2026 |
| Licence | Proprietary | Proprietary | Proprietary |
| Context window | 500K tokens | 1.05M tokens | 1.05M tokens |
| Max output | – | 66K tokens | 128K tokens |
| Input / 1M | $2.00* | $0.75* | $10.00* |
| Output / 1M | $6.00 | $3.75 | $50.00 |
| Cached input / 1M | $0.50 | $0.075 | $1.00 |
| Speed | 79 tokens/s | 249 tokens/s | 51 tokens/s |
| Input types | text, image | text, image, video, audio, pdf | text, image |
| Best for | CodingAgentsReasoning | CodingAgentsReasoningMultimodal | CodingAgentsReasoningResearch |
* Grok 4.7: Batch API is not supported for grok-4.7. Priority processing costs 2x; the US regional endpoint costs 1.1x. A faster variant (Grok 4.7 Fast, 2x price) is only in Cursor and Grok Build, not the public API.
* Gemini 3.8 Flash: Introductory price through 31 Dec 2026. From 1 Jan 2027: $1.50 input, $7.50 output, $0.15 cached input. Cache storage $0.50 per 1M tokens per hour, rising to $1.00 in 2027. Google bills no separate cache write. One flat price up to the full 1M context. Flex is also 50% off; Priority is $1.35 / $6.75 ($2.70 / $13.50 from 2027).
* GPT-6 Astra: Cache writes are 1.25x input; cache reads 0.1x; 30-minute minimum cache life. Batch and Flex: $5 in / $25 out. Fast mode (formerly Priority): 2x, $20 / $100. Ultrafast (Responses API, service_tier "ultrafast", Astra only): $60 in / $6 cached / $75 cache write / $300 out. Data residency endpoints +10%.
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- Grok 4.7
- xAI's newest flagship for coding, agentic tasks and knowledge work, with configurable reasoning effort (low to xhigh). xAI recommends it for everything except image, video and voice.
- Gemini 3.8 Flash
- Google's most intelligent Flash model, built for long-horizon software engineering, autonomous agents and complex enterprise workflows at Flash speed and cost.
- GPT-6 Astra
- OpenAI's most capable model, for the hardest end-to-end work: complex reasoning, coding, computer use, research and document creation.
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.