DeepSeek V4.1 Flash
DeepSeek · Released 10 Sep 2026 · deepseek-flash
$0.30
Input price / 1M tokens
#9 of 61, cheapest first
$1.20
Output price / 1M tokens
#8 of 61, cheapest first
1.05M
Context window
#11 of 71
209
Output tokens / second
#8 of 61
90.6%
Terminal-Bench 2.1
#1 of 17
74.2%
DeepSWE v1.1
#3 of 17
90.9%
GPQA Diamond
#14 of 27
36.8%
Humanity’s Last Exam (no tools)
#12 of 23
Summary
DeepSeek V4.1 Flash is an open-weight model from DeepSeek, released on 10 Sep 2026.
At $0.30 input and $1.20 output per million tokens, it is #9 of 61 on input price, cheapest first.
Its 1.05M-token context window ranks #11 of 71.
Measured output speed is 209 tokens per second, #8 of 61.
Smallest model in DeepSeek's new architecture family: a 552B-parameter MoE (Causal Encoder-Decoder) that activates 8B parameters per token for input and 16B for output, with native image understanding. DeepSeek says it beats V4-Pro on its benchmarks. Open weights under MIT (commercial use allowed).
Benchmarks · 6 reported
| Benchmark | Score | Bar | Rank | Source |
|---|---|---|---|---|
| Terminal-Bench 2.1DeepSeek Harness minimal mode, reasoning_effort=100 | 90.6% | #1 of 17 | DeepSeek-V4.1-Flash model cardSelf-reported | |
| DeepSWE v1.1mini-SWE harness, max effort | 74.2% | #3 of 17 | DeepSeek-V4.1-Flash model cardSelf-reported | |
| GPQA DiamondPass@1, max effort | 90.9% | #14 of 27 | DeepSeek-V4.1-Flash model cardSelf-reported | |
| Humanity’s Last Exam (no tools)no tools, full set (39.1 on text-only subset) | 36.8% | #12 of 23 | DeepSeek-V4.1-Flash model cardSelf-reported | |
| Humanity’s Last Exam (with tools)with tools | 63.9% | #4 of 16 | DeepSeek-V4.1-Flash model cardSelf-reported | |
| Codeforcesmax effort | 3,471 | Elo-style | – | DeepSeek-V4.1-Flash model cardSelf-reported |
Self-reported means the lab ran the test itself. A rank appears only where several labs report the same version of a benchmark.
Similar models
Models with the most benchmarks in common with DeepSeek V4.1 Flash, and the closest scores.
Terminal-Bench 2.1
#1 of 17Terminal tasks, earlier set
- DeepSeek V4.1 Flash90.6
- DeepSeek V4 Pro87.9
- Kimi K388.3
- Qwen3.8-Max86.6
- GLM-5.281
- GPT-5.6 Sol88.8
DeepSWE v1.1
#3 of 17Long coding tasks in real repos
- DeepSeek V4.1 Flash74.2
- DeepSeek V4 Pro62.7
- Kimi K367.5
- GPT-5.6 Sol72.7
GPQA Diamond
#14 of 27Graduate-level science questions
- DeepSeek V4.1 Flash90.9
- DeepSeek V4 Pro92.4
- Kimi K393.5
- Qwen3.8-Max92.6
- GLM-5.291.2
- GPT-5.6 Sol94.6
Humanity’s Last Exam (no tools)
#12 of 23Expert questions, no search or code
- DeepSeek V4.1 Flash36.8
- DeepSeek V4 Pro42.7
- Kimi K343.5
- Qwen3.8-Max43.6
- GLM-5.240.5
Compare DeepSeek V4.1 Flash side by side
Pre-filled with the two closest models. Swap any of them.
Sources
- price, peak/off-peak, context, max output, vision, api id: api-docs.deepseek.com/quick_start/pricing, checked 2 Oct 2026
- release date, architecture, price effective date, V4-Flash retirement: api-docs.deepseek.com/news/news260910, checked 2 Oct 2026
- license, parameters, benchmarks: huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash, checked 2 Oct 2026
- output speed: artificialanalysis.ai/models/deepseek-v4-1-flash, checked 2 Oct 2026
Building on DeepSeek V4.1 Flash?
Get the architecture right before the bill arrives.
We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.