Skip to main content
Back to News
Technology
2 min read
CN

DeepSeek V4 Pro Deep-Dive Finds Post-Training Gains, Extrapolated 1M Context, Kimi K3 Parity

The AMW Read

Independent testing reveals DeepSeek V4 Pro's agentic gains stem from post-training rather than pretraining scale, mirroring GLM-5.3's identical strategy, and shows its 1M-context claim rests on YaRN extrapolation, not native training.
NoveltySignificance
Foundation Models · Case StudiesScaling Laws
DeepSeek AI
DeepSeek AI

Foundation Models / LLMs

View Company Profile

DeepSeek V4 Pro Deep-Dive Finds Post-Training Gains, Extrapolated 1M Context, Kimi K3 Parity

An independent review by Chinese outlet Leiphone tested DeepSeek's V4 Pro model and its open-source Harness agent framework — benchmark probes, seven coding-task runs, a 920,000-token long-context test, and a pass through roughly 198,000 lines of released code. The "0813" V4 Pro release showed large agentic-benchmark jumps over the earlier Pro-Preview (Terminal Bench 2.1: 72.1 to 87.9; DeepSWE: 12.8 to 62.7), even though DeepSeek never confirmed a full post-training redo — official notes describe it as Preview's architecture plus an added speculative-decoding module. The review also found V4 Pro's advertised 1-million-token context window comes from YaRN extrapolation of a native 64K window, not training at that length.

That matters because DeepSeek is a closely watched frontier lab, and the finding suggests its cost-efficiency reputation now leans on post-training tuning rather than fresh pretraining scale. On DeepSeek's own benchmark table, Kimi K3 still beats V4 Pro-0813 on most agentic tasks (Terminal Bench 2.1: 88.3 vs 87.9; DeepSWE: 67.5 vs 62.7), and GLM-5.3, released a day apart, reused the same 744B-parameter base with no re-pretraining, gaining entirely from extended post-training. DeepSeek's own Harness framework defaults to Flash, not Pro. Per the AI Market Watch index, name-matched DeepSeek coverage rose to 105 items in the last 90 days from 82 the prior quarter (pipeline-ingested sources only), consistent with intensifying scrutiny of every pricing and technical claim since the V4 Flash launch and price hikes.

For builders, sticker price is the wrong comparison: V4 Pro's $1.32-per-million-token input rate sits well above the $0.33 open-model median, yet its measured per-task cost ranked second-lowest of 106 open models tested, because DeepSeek engineers its request format — keeping the full tool catalog live even in read-only "plan" mode — to hold KV-cache hit rates above 90%. For investors tracking DeepSeek's reported $7.4B raise at a $74B valuation, an extrapolated rather than natively trained 1M-context window is a concrete diligence point against claims of parity with closed frontier models.

#DeepSeek #FoundationModels #OpenWeightAI #AgenticBenchmarks #LLMPricing #KimiK3

#DeepSeek V4 Pro#Harness framework#Kimi K3#GLM-5.3#post-training scaling#KV-cache economics

How This Connects

Based on Foundation Models · Case Studies

  1. 21h agoOpenAI's GPT-6 Astra reportedly runs recurrent-depth transformers — Alibaba-affiliated research already diagnosed the catch.OpenAI
  2. 1d agoNvidia is reportedly weighing a $2.5 billion investment in Thinking Machines Lab that would value the startup at roughly $40 billion.Thinking Machines Lab
  3. 5d agoDeepSeek V4 Pro Deep-Dive Finds Post-Training Gains, Extrapolated 1M Context, Kimi K3 Parity · THIS ARTICLE
  4. 5d agoAnthropic Signs Reported $35B Lambda Cloud Deal for Texas AI ComputeAnthropic
  5. 3w agoDeepSeek V4 Pro (DeepSeek-V4-Pro-0813) has gone live on the company's API, bringing enhanced agent c...DeepSeek V4 Pro
  6. 1mo agoMoonshot AI launches Kimi K3, a 2.8 trillion-parameter open-weight model, claiming performance near US frontier labsMoonshot AI

Related News

More news from DeepSeek AI

Stay updated with the latest news and announcements from DeepSeek AI.

View all DeepSeek AI news

Discover AI Startups

Explore 5,000+ AI companies with VC-grade analysis, funding data, and investment insights.

Explore Dashboard