ByteDance (字节跳动) is reportedly investing up to RMB 10 trillion (over $1 trillion) in its Seed AI div...
The AMW Read
The article updates ByteDance's position in the Chinese foundation-model race with a strategic investment and reveals a new frontier-scale model plan, but doesn't overturn an existing case study.
ByteDance (字节跳动) is reportedly investing up to RMB 10 trillion (over $1 trillion) in its Seed AI division, explicitly rejecting the distillation shortcut in favor of building frontier models from scratch. In a recent all-hands meeting, CEO Zhang Yiming (张一鸣) declared that ByteDance should sacrifice short-term gains for long-term goals, refusing to distill outputs from rival labs. This stance has already cost ByteDance in the rankings, with its Seed 2.1 Pro falling to 21st place on Artificial Analysis, far behind Chinese rivals like Kimi K3 Max (2nd) and Qwen 3.8 Max (4th). The company is instead advancing a plan to train a 10-trillion-parameter model, led by Xiang Liang (项亮), head of Seed Foundation, with pre-training data lead Shen Ke (沈科) — a scale several times larger than Alibaba's Qwen 3.8-Max (2.4T) or Moonshot's K3 (2.8T).
ByteDance's bet is a direct counter-narrative to the industry-wide reliance on distillation, especially among Chinese labs that have used OpenAI and Anthropic outputs to leapfrog in benchmarks. By refusing to distill, ByteDance is betting on proprietary data — its Douyin and TikTok feed generates over 80 million videos daily — and on building a full-stack engineering infrastructure to support frontier-scale training. This is a high-risk, capital-intensive strategy that could either cement ByteDance's position as a front-runner in the next generation of AI or leave it trailing as peers like DeepSeek and Alibaba iterate faster on existing architectures. The 10-trillion-parameter ambition underscores a belief that scaling laws still hold and that size, not speed, will ultimately decide the frontier.
For builders and investors, this signals two things. First, expect a growing bifurcation in AI strategy: distillation-based fast-followers will ship incremental model updates, while from-scratch players will make slower but potentially more foundational leaps, with ByteDance's move possibly triggering a similar shift among other major labs. Second, the sheer scale of capital and engineering required — ByteDance's 2026 capex for AI infrastructure has already been raised to over RMB 200 billion (around $28 billion), contributing to a reported 70% profit drop — means the compute and data moat is deepening. For those building on top of models, this could mean more differentiated and capable foundations, but also a faster consolidation of the frontier-player pool. Investors should watch whether ByteDance's from-scratch approach can close the benchmark gap before the company's patience runs out, as a failure would not only be a financial setback but also a validation of distillation as a legitimate path.

