DeepSeek Launches V4.1-Flash at $0.003 per Million Off-Peak Cached Tokens, Claims Benchmark Edge Over GPT-5.6 and Claude Opus 5
The AMW Read
Extends DeepSeek's known pricing-and-benchmark arc rather than resolving it, but self-reported wins over top-tier rivals carry segment-wide pricing-competition significance.
DeepSeek Launches V4.1-Flash at $0.003 per Million Off-Peak Cached Tokens, Claims Benchmark Edge Over GPT-5.6 and Claude Opus 5
DeepSeek has released V4.1-Flash, a new model priced at $0.003 per million tokens for off-peak cached input. The company says its benchmark results surpass OpenAI's GPT-5.6 and Anthropic's Claude Opus 5, positioning V4.1-Flash as the latest entrant in the price-performance race among frontier model labs.
The launch complicates DeepSeek's own pricing narrative. In mid-August, the company raised API prices as much as 12x and added peak-time pricing, citing compute constraints as it shifted from aggressive discounting toward sustainable revenue. An off-peak cached rate this low reads as demand segmentation, not a reversal — pushing latency-tolerant workloads into cheap windows while protecting margin at peak hours. It also arrives with a caution flag from DeepSeek's own record: independent testing of the prior V4 Flash found its agentic gains traced to post-training, and the model failed 46.2% of real-world agent tasks despite topping leaderboards. Self-reported wins over GPT-5.6 and Claude Opus 5 deserve the same scrutiny.
For builders, test the off-peak tier on production tasks rather than trusting leaderboard rank, given the benchmark-to-reality gap DeepSeek's own V4 Flash exposed. For investors, the timing matters: per the AI Market Watch index, DeepSeek remains privately funded by parent High-Flyer Quant with no outside VC capital as of July 2025 and a first external round in negotiation as of May 2026 (index coverage, not a full census) — a cheap, headline model launch supports a valuation story more than it proves a margin one.

