Skip to content
Free tool · Updated 2 Oct 2026

DeepSeek V4 Pro vs Kimi K3 vs DeepSeek V4.1 Flash

Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.

DeepSeek V4 ProKimi K3DeepSeek V4.1 Flash

Cheapest blended

DeepSeek V4.1 Flash

$0.525 / 1M tokens

Largest context

DeepSeek V4 ProKimi K3DeepSeek V4.1 Flash

Tie at 1.05M tokens

Fastest output

DeepSeek V4.1 Flash

209 tokens/s

Top on Terminal-Bench 2.1

DeepSeek V4.1 Flash

90.6%

Specifications

SpecDeepSeek V4 ProDeepSeekKimi K3MoonshotDeepSeek V4.1 FlashDeepSeek
ProviderDeepSeekMoonshotDeepSeek
Released24 Apr 2026Jul 202610 Sep 2026
LicenceOpen weightOpen weightOpen weight
Context window1.05M tokens1.05M tokens1.05M tokens
Max output393K tokens1.05M tokens393K tokens
Input / 1M$1.32*$3.00*$0.30*
Output / 1M$3.96$15.00$1.20
Cached input / 1M$0.044$0.30$0.006
Speed104 tokens/s34 tokens/s209 tokens/s
Input typestexttext, image, videotext, image
Best forCodingAgentsReasoningLong contextCodingAgentsReasoningLong contextCodingAgentsBudgetVision

* DeepSeek V4 Pro: Peak price shown. Off-peak is half: $0.66 in, $1.98 out, $0.022 cache hit. Peak hours are 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, excluding Chinese public holidays; weekends are off-peak all day. This peak/off-peak pricing started 16 Aug 2026 16:00 UTC and replaced the flat $0.435/$0.87 price that DeepSeek made permanent on 22 May 2026 (launch list price was $1.74/$3.48).

* Kimi K3: Cache write $3.00 (5-minute TTL, default) or $6.00 (1-hour TTL); cache hits $0.30. Prices exclude taxes.

* DeepSeek V4.1 Flash: Peak price shown. Off-peak is half: $0.15 in, $0.60 out, $0.003 cache hit. Peak hours 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, excluding Chinese public holidays. Price took effect 10 Sep 2026 04:00 UTC. Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are routed here and billed at this price.

Benchmark scores

Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.

DeepSeek V4 ProKimi K3DeepSeek V4.1 Flash
Terminal-Bench 2.187.9%88.3%90.6%
Humanity’s Last Exam (no tools)42.7%43.5%36.8%
Humanity’s Last Exam (with tools)60%56%63.9%
DeepSWE v1.162.7%67.5%74.2%
GPQA Diamond92.4%93.5%90.9%

Price per 1M tokens

DeepSeek V4 ProKimi K3DeepSeek V4.1 Flash
Input$1.32$3.00$0.30
Output$3.96$15.00$1.20
Cached input$0.044$0.30$0.006

Context window and speed

DeepSeek V4 ProKimi K3DeepSeek V4.1 Flash
Context1.05M1.05M1.05M
Tokens/s10434209

Notes

DeepSeek V4 Pro
DeepSeek's flagship MoE model, 1.6T total and 49B active parameters, open weights under MIT (commercial use allowed, no extra conditions). The API now serves the GA build DeepSeek-V4-Pro-0813 (13 Aug 2026), which DeepSeek says greatly improves agent performance and adds low/high/max reasoning effort.
Kimi K3
Moonshot's flagship and the first open 3T-class model: 2.8T total, 104B active parameters, native vision, 1M context. Kimi K3 License allows commercial use, but Model-as-a-Service operators with over $20M revenue in 12 months need a separate agreement, and products over 100M MAU or $20M monthly revenue must display "Kimi K3".
DeepSeek V4.1 Flash
Smallest model in DeepSeek's new architecture family: a 552B-parameter MoE (Causal Encoder-Decoder) that activates 8B parameters per token for input and 16B for output, with native image understanding. DeepSeek says it beats V4-Pro on its benchmarks. Open weights under MIT (commercial use allowed).

Still torn between two?

Test them on your own prompts.

Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.