DeepSeek V4 Pro Deep-Dive Finds Post-Training Gains, Extrapolated 1M Context, Kimi K3 Parity
The AMW Read
Independent testing reveals DeepSeek V4 Pro's agentic gains stem from post-training rather than pretraining scale, mirroring GLM-5.3's identical strategy, and shows its 1M-context claim rests on YaRN extrapolation, not native training.
DeepSeek V4 Pro Deep-Dive Finds Post-Training Gains, Extrapolated 1M Context, Kimi K3 Parity
An independent review by Chinese outlet Leiphone tested DeepSeek's V4 Pro model and its open-source Harness agent framework — benchmark probes, seven coding-task runs, a 920,000-token long-context test, and a pass through roughly 198,000 lines of released code. The "0813" V4 Pro release showed large agentic-benchmark jumps over the earlier Pro-Preview (Terminal Bench 2.1: 72.1 to 87.9; DeepSWE: 12.8 to 62.7), even though DeepSeek never confirmed a full post-training redo — official notes describe it as Preview's architecture plus an added speculative-decoding module. The review also found V4 Pro's advertised 1-million-token context window comes from YaRN extrapolation of a native 64K window, not training at that length.
That matters because DeepSeek is a closely watched frontier lab, and the finding suggests its cost-efficiency reputation now leans on post-training tuning rather than fresh pretraining scale. On DeepSeek's own benchmark table, Kimi K3 still beats V4 Pro-0813 on most agentic tasks (Terminal Bench 2.1: 88.3 vs 87.9; DeepSWE: 67.5 vs 62.7), and GLM-5.3, released a day apart, reused the same 744B-parameter base with no re-pretraining, gaining entirely from extended post-training. DeepSeek's own Harness framework defaults to Flash, not Pro. Per the AI Market Watch index, name-matched DeepSeek coverage rose to 105 items in the last 90 days from 82 the prior quarter (pipeline-ingested sources only), consistent with intensifying scrutiny of every pricing and technical claim since the V4 Flash launch and price hikes.
For builders, sticker price is the wrong comparison: V4 Pro's $1.32-per-million-token input rate sits well above the $0.33 open-model median, yet its measured per-task cost ranked second-lowest of 106 open models tested, because DeepSeek engineers its request format — keeping the full tool catalog live even in read-only "plan" mode — to hold KV-cache hit rates above 90%. For investors tracking DeepSeek's reported $7.4B raise at a $74B valuation, an extrapolated rather than natively trained 1M-context window is a concrete diligence point against claims of parity with closed frontier models.
#DeepSeek #FoundationModels #OpenWeightAI #AgenticBenchmarks #LLMPricing #KimiK3

