Technology

DeepSeek V4.1 Flash vs V4 Flash (2026): One Week Apart, a Full Price Table Apart

DeepSeek V4.1 Flash vs V4 Flash in 2026: price tables, native vision, Pro redirect, and what changes for cached workloads - with a clear decision rule.

5
V4.1 Flash — current stable API name, native multimodality, lower price table
vs
0
V4 Flash — temporary August build, retired and redirecting to deepseek-flash
Quick Verdict

V4.1 Flash is the only current entry of the Flash line: same context, roughly a third of the price, native vision in one ID, stable name. V4 Flash survives as the August price table behind the redirect and as a historical snapshot for reproducible evaluations. Decision rule: count the tiers separately (peak 01:00–04:00 and 06:00–10:00 UTC, otherwise half price), let the deepseek-flash pin do its job, and budget on the new table — the old Pro assumptions are no longer true since 14.09.2026.

Detailed Comparison

A side-by-side analysis of key factors to help you make the right choice.

Factor
V4.1 Flash — current stable API name, native multimodality, lower price tableRecommended
V4 Flash — temporary August build, retired and redirecting to deepseek-flashWinner
Input Price
Cache-hit $0.003, cache-miss $0.15 off-peak (double at peak)
Cache-hit $0.007, cache-miss $0.22 off-peak (double at peak)
Output Price
$0.60 per 1M tokens off-peak, $1.20 at peak
$0.66 off-peak, $1.32 at peak
Multimodality
Native image understanding inside the single deepseek-flash ID
Image support only via separate deepseek-v4-flash-vision-exp ID (now retired)
Context Window
1M-token context, 384K max output
1M-token context, 384K max output
Api Name Stability
Stable, official API name in the current price table
Temporary build, replaced on 10.09.2026 — name no longer resolves on its own
Pro Plan Effect
From 14.09.2026 every deepseek-v4-pro request is billed at the V4.1-Flash table (about 77% cheaper cache-miss, 70% cheaper output)
No separate effect — the old Pro price table is fully superseded
Peak Offpeak Tiering
Off-peak is exactly half of peak; peak = Mon-Fri 01:00-04:00 and 06:00-10:00 UTC
Same two-tier logic, just at the higher August prices
Local Deployment
Open MIT weights, 552B MoE with 8B active in / 16B out, MXFP4, about 510 GB full stack
Same architecture family and licence; smaller weight snapshot on disk
Total Score5/ 80/ 83 ties
Input Price
V4.1 Flash — current stable API name, native multimodality, lower price table
Cache-hit $0.003, cache-miss $0.15 off-peak (double at peak)
V4 Flash — temporary August build, retired and redirecting to deepseek-flash
Cache-hit $0.007, cache-miss $0.22 off-peak (double at peak)
Output Price
V4.1 Flash — current stable API name, native multimodality, lower price table
$0.60 per 1M tokens off-peak, $1.20 at peak
V4 Flash — temporary August build, retired and redirecting to deepseek-flash
$0.66 off-peak, $1.32 at peak
Multimodality
V4.1 Flash — current stable API name, native multimodality, lower price table
Native image understanding inside the single deepseek-flash ID
V4 Flash — temporary August build, retired and redirecting to deepseek-flash
Image support only via separate deepseek-v4-flash-vision-exp ID (now retired)
Context Window
V4.1 Flash — current stable API name, native multimodality, lower price table
1M-token context, 384K max output
V4 Flash — temporary August build, retired and redirecting to deepseek-flash
1M-token context, 384K max output
Api Name Stability
V4.1 Flash — current stable API name, native multimodality, lower price table
Stable, official API name in the current price table
V4 Flash — temporary August build, retired and redirecting to deepseek-flash
Temporary build, replaced on 10.09.2026 — name no longer resolves on its own
Pro Plan Effect
V4.1 Flash — current stable API name, native multimodality, lower price table
From 14.09.2026 every deepseek-v4-pro request is billed at the V4.1-Flash table (about 77% cheaper cache-miss, 70% cheaper output)
V4 Flash — temporary August build, retired and redirecting to deepseek-flash
No separate effect — the old Pro price table is fully superseded
Peak Offpeak Tiering
V4.1 Flash — current stable API name, native multimodality, lower price table
Off-peak is exactly half of peak; peak = Mon-Fri 01:00-04:00 and 06:00-10:00 UTC
V4 Flash — temporary August build, retired and redirecting to deepseek-flash
Same two-tier logic, just at the higher August prices
Local Deployment
V4.1 Flash — current stable API name, native multimodality, lower price table
Open MIT weights, 552B MoE with 8B active in / 16B out, MXFP4, about 510 GB full stack
V4 Flash — temporary August build, retired and redirecting to deepseek-flash
Same architecture family and licence; smaller weight snapshot on disk

Key Statistics

Real data from verified industry sources to support your decision.

Cache-miss input after Pro redirect: $0.66 -> $0.15 per 1M tokens (about 77% cheaper)

DeepSeek changelog

Output price after Pro redirect: $1.98 -> $0.60 per 1M tokens (about 70% cheaper)

DeepSeek changelog

Sample task (10K cache-miss in + 2K out) at peak: about $0.0106 (August) -> about $0.0030 (from 10.09.)

Apidog price overview

Combined saving over one peak + one off-peak 1M batch: about $17 per million tokens

Apidog price overview

Throughput: 221-235 tok/s official, up to 400 tok/s in third-party runs

Coursiv / WorldofAI test log

Effective dates: new prices 10.09.2026 04:00 UTC; Pro redirect 14.09.2026 04:00 UTC

DeepSeek changelog

All statistics come from verified third-party sources. Source, year, and direct link are shown on each metric.

When to Choose Each Option

Clear guidance based on your specific situation and needs.

Choose V4.1 Flash — current stable API name, native multimodality, lower price table when...

  • You run production agents on the paid API and forecast costs by the minute.
  • You need image understanding and text in a single, stable model ID.
  • You have cached workloads - the cache-hit rate now applies at $0.003 per 1M.
  • You want the September 14 Pro redirect to do its work without touching code.

Choose V4 Flash — temporary August build, retired and redirecting to deepseek-flash when...

  • You reproduce an August 2026 evaluation and need the old price snapshot for it.
  • You run the older local weight snapshot with the smaller disk footprint.
  • You document the price history of the Flash line as its own step.
  • Your stack still resolves the retired IDs and you want the redirect behaviour verified.

Our Recommendation

V4.1 Flash is the only current entry of the Flash line: same context, roughly a third of the price, native vision in one ID, stable name. V4 Flash survives as the August price table behind the redirect and as a historical snapshot for reproducible evaluations. Decision rule: count the tiers separately (peak 01:00–04:00 and 06:00–10:00 UTC, otherwise half price), let the deepseek-flash pin do its job, and budget on the new table — the old Pro assumptions are no longer true since 14.09.2026.

Frequently Asked Questions

Common questions about this comparison answered.

Same context size (1M tokens, 384K max output), but V4.1 Flash brings a lower price table (cache-miss $0.15, output $0.60 off-peak), native image understanding in one ID, and the status of a stable, official model name. V4 Flash was the temporary August build and is now retired behind a redirect.
With the V4.1-Flash release on 10.09.2026 at 04:00 UTC, deepseek-flash became the model name in the API. From 14.09.2026 (04:00 UTC), every request to deepseek-v4-pro is routed automatically to the V4.1-Flash price.
Peak is Monday to Friday, 01:00-04:00 and 06:00-10:00 UTC. All other hours are off-peak and cost exactly half. The simple lever: move batch jobs into the off-peak windows and keep only latency-critical calls at peak.
Mostly not. Pinned deepseek-flash captures the savings automatically. Two checks are worth it: whether old Pro cost assumptions are still hard-coded, and whether your model list still names deepseek-v4-flash or deepseek-v4-flash-vision-exp - both IDs now redirect to deepseek-flash, so hardcoded lists show one line less than expected.

Need help deciding?

Book a free 30-minute consultation and we'll help you determine the best approach for your specific project.

Free consultation
No obligation
Response within 24h