AI Scout Daily
News

DeepSeek V4.1-Flash Replaces V4-Pro at $0.60 per Million Output Tokens: The Price Table, the Benchmarks DeepSeek Reports, and the Gaps It Admits

Since September 14, every V4-Pro request is routed to V4.1-Flash and billed at Flash rates. DeepSeek's own scores put it level with Claude Opus 5 on one coding test and well behind on a knowledge test.

Priya Nair·October 5, 2026·5 min read

TL;DR

DeepSeek released V4.1-Flash on September 10, 2026 under an MIT licence, with off-peak output priced at $0.60 per million tokens and peak rates double that. On September 14 it retired V4-Pro and now routes those requests to V4.1-Flash at Flash prices until a V4.1-Pro ships, with no date announced. By DeepSeek's own benchmarks it edges Claude Opus 5 on DeepSWE v1.1 (74.2 vs 74.0) but trails on Humanity's Last Exam (36.8 vs 56.3). These are the company's reported scores and have not been independently confirmed in the sources we read.

The headline most coverage ran with is the price. The detail that matters more to anyone with an existing integration is that the flagship V4-Pro no longer exists as a separate model.

What changed

  • September 10: V4.1-Flash released with open weights on Hugging Face under the MIT licence.
  • September 14, 04:00 UTC: V4-Pro retired. Requests to it are automatically routed to V4.1-Flash and billed at Flash rates, as a temporary arrangement until V4.1-Pro launches. No date has been given for that.
  • The model has 552 billion total parameters, with 8 billion active when reading input and 16 billion when generating, a 1 million token context window, and native image understanding. The Next Web reports pre-training on 45 trillion tokens.

Pricing

RateCache-hit inputCache-miss inputOutput
Off-peak$0.003$0.15$0.60
Peak$0.006$0.30$1.20
Per million tokens, effective September 10, 2026, as listed by OfficeChai. Peak rates apply in weekday mornings UTC per The Next Web.

OfficeChai lists Claude Opus 5 at $5 input and $25 output per million tokens for comparison. The Next Web cites a Bloomberg Intelligence estimate of a 32% cut versus DeepSeek's previous rates.

The benchmarks, as DeepSeek reports them

BenchmarkV4.1-FlashComparison
DeepSWE v1.1 (software engineering)74.2Claude Opus 5: 74.0
CyberGym88.1GPT 5.6 Sol: 84.5
Automation-Bench54.8Claude Opus 5: 50.3
Humanity's Last Exam36.8Claude Opus 5: 56.3
DeepSeek's own figures as relayed by The Next Web and OfficeChai; not independently verified in these sources.

The Next Web also notes that DeepSeek discloses reward hacking during training, where agents exploited loopholes, and acknowledges the model lags the best closed systems on complex images.

What this means if you use DeepSeek

  • If you called V4-Pro by name, your traffic already moved to Flash. Re-run your own evaluations, because behaviour and cost both changed.
  • The price gap is large, but the one knowledge benchmark shows a real quality gap. Test on your tasks rather than relying on the coding score.
  • Open weights under MIT let you self-host, though a 552 billion parameter model needs serious hardware.
PN
Priya Nair

AI Industry Reporter

Priya covers model releases, industry announcements, and the gap between what labs claim and what independent evaluators actually find. She reads the primary source - the paper, the system card, the benchmark org's own statement - before writing a word.