DeepSeek V4.1 Flash: Nine of Nine for $0.36 (and $0.13 for the Model It Replaces)
DeepSeek V4.1 Flash scored 9 out of 9 on our executed Python benchmark at a derived $0.36 per 1,000 tasks and a 15.6-second mean. The interesting part is what sits next to it in our data: DeepSeek V4-Flash, the model it supersedes, also scored 9 out of 9, at $0.13 — and it is still on the catalog with a live list price. On these nine tasks the upgrade buys 23 cents more per 1,000 tasks (0.36 − 0.13 = 0.23), 1.1 seconds more latency, and not one extra correct answer.
A point release is supposed to be the same thing, cheaper. This one is the same score, dearer, and slower — while the model it replaces is still sitting on the price list one click away.
The result
| Metric | DeepSeek V4.1 Flash |
|---|---|
| Score | 9/9 |
| Derived cost / 1,000 tasks | $0.36 |
| Mean latency | 15.6s |
| Reasoning tokens per call | 478 |
| Tokens across the suite | 834 in / 5,217 out |
| List price used for the cost | $0.15 in / $0.6 out per 1M, captured 2026-09-15 |
| List price today | $0.15 in / $0.6 out per 1M, unchanged |
| Context window | 1,048,576 |
| Measured | 2026-09-16, post-sweep |
Of the 56 models in our set that score 9 out of 9, V4.1 Flash ranks 8th on cost and 45th on latency. That is the shape of a cheap reasoning model: it is near the front of the price field and near the back of the speed field, because it thinks 478 tokens' worth before it answers and you pay for every one of them at the output rate.
V4.1 Flash against V4-Flash
| Metric | DeepSeek V4.1 Flash | DeepSeek V4-Flash |
|---|---|---|
| Score | 9/9 | 9/9 |
| Derived cost / 1,000 tasks | $0.36 | $0.13 |
| Mean latency | 15.6s | 14.5s |
| Reasoning tokens per call | 478 | 568 |
| Cost rank among the 56 at 9/9 | 8th | 5th |
| Latency rank among the 56 at 9/9 | 45th | 43rd |
| Context window | 1,048,576 | 1,048,576 |
| List price today | $0.15 in / $0.6 out per 1M | $0.076 in / $0.15 out per 1M |
| Sweep | post-sweep, run 2026-09-16 | core-13, no run date stored |
The older model wins every column that touches your bill or your clock. It is cheaper, it is marginally faster, it carries the same million-token context, and it clears the same nine tasks. The one place V4.1 Flash is ahead is internal economy: 90 fewer reasoning tokens per call (568 − 478 = 90). That did not translate into a cheaper run, because the price per token moved in the other direction. On list prices as they stand today, V4.1 Flash's output rate is exactly four times its predecessor's: $0.6 ÷ $0.15 = 4.
None of which means V4.1 Flash is a worse model. It means these nine tasks cannot tell the difference. Nine self-contained Python functions — parsing, intervals, a token bucket — cannot separate a frontier model from a competent small one, and they certainly cannot separate two siblings from the same lab. A ceiling of 9/9 is a ceiling: everything above “solves all nine” is invisible to us. What the suite can price precisely is the cost of work this size, and at this size the newer model is the more expensive way to get an identical answer.
The two dollar figures are not the same age
Here is the caveat we would want someone to hand us. The $0.36 was derived from tokens measured on 2026-09-16 at the list price we captured on 2026-09-15. The $0.13 was derived at the list price in force on 2026-07-17 — and that entry does not store the price it used, because it predates the field. It also does not store its own token counts.
So we cannot re-derive $0.13 at today's $0.076 / $0.15, and we will not pretend otherwise. Two months of list-price movement sit between those two numbers, and DeepSeek's flash rate has moved more than once in that window — our original V4-Flash review caught it mid-move. The honest reading of this pair is: two models, same score, one measured cheap in July and one measured dearer in September, on a lane whose prices do not sit still. The direction of the gap is solid. The exact 23 cents is a July-to-September comparison, not a same-day one.
Where $0.36 sits on the ladder
The band this chart covers is narrow in dollars and enormous in ratio. Ling 3.0 Flash VL holds the floor at $0.07, derived 2026-09-15, with a 4.4-second mean; GPT-5.4 Mini is the fastest 9/9 we have at 2.3 seconds and still costs $0.53, derived at its 2026-07-30 price. V4.1 Flash is in the middle of that pack on price and near the bottom on speed, which makes it a hard sell against its own neighbours: $0.36 ÷ $0.08 = 4.5, exactly, against DeepSeek V3.2 at its 2026-07-30 price, which does the same nine tasks at 7.1 seconds with zero reasoning tokens.
Look up instead of sideways and the picture flips. GPT-6 Astra scores the same 9 out of 9 at $8.19, derived at its 2026-09-15 price — $8.19 ÷ $0.36 = 22.75, exactly — and the priciest 9/9 in the whole set, Sakana Fugu Ultra V2 at $57.06 on the same 2026-09-15 capture, is $57.06 ÷ $0.36 = 158.5 times V4.1 Flash. Against the top of the market, 36 cents is a rounding error. Against the bottom of the market, it is four and a half times V3.2's derived cost for the same nine answers. Which comparison matters depends entirely on whether your workload looks like our nine tasks — and if it does, you should be shopping in the cheapest tier, not here.
One neighbour we deliberately left off the chart: IBM Granite 4.2 8B came in at a derived $0.43 at its 2026-09-15 price, which would place it right beside V4.1 Flash. We publish no score for it. Two of its nine tasks failed at the API layer after retries, leaving 7 of 9 scored, and a 7-of-7 on a run that never finished is not comparable to a 9/9 on a run that did. It is marked excluded in our data and it stays out of every ranking.
Three ids in one lane
DeepSeek's flash lane is easy to mis-pin. In the alias snapshot we captured on 2026-09-15, deepseek/deepseek-flash-latest resolves to deepseek/deepseek-v4.1-flash — the $0.36 model. Meanwhile deepseek/deepseek-v4-flash-latest resolves to deepseek/deepseek-v4-flash-0731, a dated snapshot that is a different entry again from the plain deepseek/deepseek-v4-flash we scored at $0.13. Three ids, one family, and the one that sounds most like “the current flash model” is the dearest of them.
That dated snapshot carries one more trap. Of the 85 base models with a :batch endpoint in our 2026-09-15 capture, 70 price batch output at exactly half the standard rate. deepseek/deepseek-v4-flash-0731 is one of the 15 that do not: standard $0.06 / $0.11, batch $0.11 / $0.33 — batch output at 300% of the standard rate, the second-worst ratio in the whole capture. If you reach for :batch on that id expecting a discount, you triple your output bill instead. Pinning an exact id is the cheap habit here; so is checking the batch endpoint before you route to it.
If you are choosing inside the family rather than across it, our Pro-versus-Flash piece covers the tier above, and the DeepSeek pricing page tracks the rates.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why Granite is excluded rather than scored low. Cost is derived from measured token counts at the list price captured on the stated date — it is not a billing statement. Runs go through OpenRouter. Full method on the methodology page, and the broader cost comparison lives in the coding cost benchmark.
What we did not measure
- Anything above the ceiling. V4.1 Flash and V4-Flash both scored 9/9, so our suite has no resolution left to distinguish them. If the point release is better at long refactors, agent loops, or debugging, we did not see it and would not have.
- V4-Flash at today's prices. Its entry stores neither the price it was costed at nor its token counts, so its $0.13 cannot be repriced against the current $0.076 / $0.15. We show the figure as measured and date it.
- The 1,048,576-token context. Our prompts are a few hundred tokens. Both models claim a million and we exercised none of it.
- Same-day comparison. V4-Flash is a core-13 entry with no stored run date; V4.1 Flash ran 2026-09-16 in the post-sweep. Different days, possibly different serving conditions.
- Repeat runs. One scored attempt per task, no averaging. A 1.1-second latency gap between two models is inside the noise of a single run and we would not defend it.
- The batch endpoint in practice. We read the 300% ratio off the price capture; we did not send traffic through
:batchto confirm what it charges.
For the wider cheap end of the field, the cheap coding roundup and the GLM 5.3 Flash review ($0.34, derived 2026-08-31) cover the models directly on either side of this one.
FAQ
How much does DeepSeek V4.1 Flash cost? $0.15 per million input tokens and $0.6 output, as captured 2026-09-15 and unchanged today. On our suite that derived to $0.36 per 1,000 tasks.
Is DeepSeek V4.1 Flash better than V4-Flash? Not on our nine tasks — both scored 9 out of 9. V4-Flash did it at $0.13 and 14.5 seconds against $0.36 and 15.6 seconds. Anything the point release improved is above our ceiling.
Is V4-Flash still available? Yes, with a live list price of $0.076 in / $0.15 out per million. It is a separate id from deepseek-v4-flash-0731, which is what the deepseek-v4-flash-latest alias points at.
Is it fast? No. 15.6 seconds mean, 45th of the 56 models that scored 9 out of 9. The fastest 9/9 in the set answers in 2.3 seconds.
What is the cheapest model that clears the suite? Ling 3.0 Flash VL at $0.07 per 1,000 tasks, derived 2026-09-15, with DeepSeek V3.2 at $0.08, derived 2026-07-30, behind it.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab