Skip to content
Free tool · Updated 2 Oct 2026

DeepSeek V4.1 Flash vs DeepSeek V4 Pro vs Kimi K3

Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.

DeepSeek V4.1 FlashDeepSeek V4 ProKimi K3

Cheapest blended

DeepSeek V4.1 Flash

$0.525 / 1M tokens

Largest context

DeepSeek V4.1 FlashDeepSeek V4 ProKimi K3

Tie at 1.05M tokens

Fastest output

DeepSeek V4.1 Flash

209 tokens/s

Top on Terminal-Bench 2.1

DeepSeek V4.1 Flash

90.6%

Specifications

SpecDeepSeek V4.1 FlashDeepSeekDeepSeek V4 ProDeepSeekKimi K3Moonshot
ProviderDeepSeekDeepSeekMoonshot
Released10 Sep 202624 Apr 2026Jul 2026
LicenceOpen weightOpen weightOpen weight
Context window1.05M tokens1.05M tokens1.05M tokens
Max output393K tokens393K tokens1.05M tokens
Input / 1M$0.30*$1.32*$3.00*
Output / 1M$1.20$3.96$15.00
Cached input / 1M$0.006$0.044$0.30
Speed209 tokens/s104 tokens/s34 tokens/s
Input typestext, imagetexttext, image, video
Best forCodingAgentsBudgetVisionCodingAgentsReasoningLong contextCodingAgentsReasoningLong context

* DeepSeek V4.1 Flash: Peak price shown. Off-peak is half: $0.15 in, $0.60 out, $0.003 cache hit. Peak hours 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, excluding Chinese public holidays. Price took effect 10 Sep 2026 04:00 UTC. Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are routed here and billed at this price.

* DeepSeek V4 Pro: Peak price shown. Off-peak is half: $0.66 in, $1.98 out, $0.022 cache hit. Peak hours are 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, excluding Chinese public holidays; weekends are off-peak all day. This peak/off-peak pricing started 16 Aug 2026 16:00 UTC and replaced the flat $0.435/$0.87 price that DeepSeek made permanent on 22 May 2026 (launch list price was $1.74/$3.48).

* Kimi K3: Cache write $3.00 (5-minute TTL, default) or $6.00 (1-hour TTL); cache hits $0.30. Prices exclude taxes.

Benchmark scores

Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.

DeepSeek V4.1 FlashDeepSeek V4 ProKimi K3
Terminal-Bench 2.190.6%87.9%88.3%
DeepSWE v1.174.2%62.7%67.5%
GPQA Diamond90.9%92.4%93.5%
Humanity’s Last Exam (no tools)36.8%42.7%43.5%
Humanity’s Last Exam (with tools)63.9%60%56%

Price per 1M tokens

DeepSeek V4.1 FlashDeepSeek V4 ProKimi K3
Input$0.30$1.32$3.00
Output$1.20$3.96$15.00
Cached input$0.006$0.044$0.30

Context window and speed

DeepSeek V4.1 FlashDeepSeek V4 ProKimi K3
Context1.05M1.05M1.05M
Tokens/s20910434

Notes

DeepSeek V4.1 Flash
Smallest model in DeepSeek's new architecture family: a 552B-parameter MoE (Causal Encoder-Decoder) that activates 8B parameters per token for input and 16B for output, with native image understanding. DeepSeek says it beats V4-Pro on its benchmarks. Open weights under MIT (commercial use allowed).
DeepSeek V4 Pro
DeepSeek's flagship MoE model, 1.6T total and 49B active parameters, open weights under MIT (commercial use allowed, no extra conditions). The API now serves the GA build DeepSeek-V4-Pro-0813 (13 Aug 2026), which DeepSeek says greatly improves agent performance and adds low/high/max reasoning effort.
Kimi K3
Moonshot's flagship and the first open 3T-class model: 2.8T total, 104B active parameters, native vision, 1M context. Kimi K3 License allows commercial use, but Model-as-a-Service operators with over $20M revenue in 12 months need a separate agreement, and products over 100M MAU or $20M monthly revenue must display "Kimi K3".

Still torn between two?

Test them on your own prompts.

Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.