DeepSeek V4 Pro vs Kimi K3 vs DeepSeek V4.1 Flash
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
DeepSeek V4.1 Flash
$0.525 / 1M tokens
Largest context
DeepSeek V4 ProKimi K3DeepSeek V4.1 Flash
Tie at 1.05M tokens
Fastest output
DeepSeek V4.1 Flash
209 tokens/s
Top on Terminal-Bench 2.1
DeepSeek V4.1 Flash
90.6%
Specifications
| Spec | DeepSeek V4 ProDeepSeek | Kimi K3Moonshot | DeepSeek V4.1 FlashDeepSeek |
|---|---|---|---|
| Provider | DeepSeek | Moonshot | DeepSeek |
| Released | 24 Apr 2026 | Jul 2026 | 10 Sep 2026 |
| Licence | Open weight | Open weight | Open weight |
| Context window | 1.05M tokens | 1.05M tokens | 1.05M tokens |
| Max output | 393K tokens | 1.05M tokens | 393K tokens |
| Input / 1M | $1.32* | $3.00* | $0.30* |
| Output / 1M | $3.96 | $15.00 | $1.20 |
| Cached input / 1M | $0.044 | $0.30 | $0.006 |
| Speed | 104 tokens/s | 34 tokens/s | 209 tokens/s |
| Input types | text | text, image, video | text, image |
| Best for | CodingAgentsReasoningLong context | CodingAgentsReasoningLong context | CodingAgentsBudgetVision |
* DeepSeek V4 Pro: Peak price shown. Off-peak is half: $0.66 in, $1.98 out, $0.022 cache hit. Peak hours are 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, excluding Chinese public holidays; weekends are off-peak all day. This peak/off-peak pricing started 16 Aug 2026 16:00 UTC and replaced the flat $0.435/$0.87 price that DeepSeek made permanent on 22 May 2026 (launch list price was $1.74/$3.48).
* Kimi K3: Cache write $3.00 (5-minute TTL, default) or $6.00 (1-hour TTL); cache hits $0.30. Prices exclude taxes.
* DeepSeek V4.1 Flash: Peak price shown. Off-peak is half: $0.15 in, $0.60 out, $0.003 cache hit. Peak hours 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, excluding Chinese public holidays. Price took effect 10 Sep 2026 04:00 UTC. Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are routed here and billed at this price.
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- DeepSeek V4 Pro
- DeepSeek's flagship MoE model, 1.6T total and 49B active parameters, open weights under MIT (commercial use allowed, no extra conditions). The API now serves the GA build DeepSeek-V4-Pro-0813 (13 Aug 2026), which DeepSeek says greatly improves agent performance and adds low/high/max reasoning effort.
- Kimi K3
- Moonshot's flagship and the first open 3T-class model: 2.8T total, 104B active parameters, native vision, 1M context. Kimi K3 License allows commercial use, but Model-as-a-Service operators with over $20M revenue in 12 months need a separate agreement, and products over 100M MAU or $20M monthly revenue must display "Kimi K3".
- DeepSeek V4.1 Flash
- Smallest model in DeepSeek's new architecture family: a 552B-parameter MoE (Causal Encoder-Decoder) that activates 8B parameters per token for input and 16B for output, with native image understanding. DeepSeek says it beats V4-Pro on its benchmarks. Open weights under MIT (commercial use allowed).
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.