Skip to main content
Back to News
Moonshot AI's Kimi K3 hits 709 tokens/sec on 16 Google TPU v7 chips, beating GB200's 452 in a like-f...
Technology
3 min read
CN

Moonshot AI's Kimi K3 hits 709 tokens/sec on 16 Google TPU v7 chips, beating GB200's 452 in a like-f...

The AMW Read

First credible benchmark showing TPU inference beating Nvidia on identical models and engines via software, challenging the assumption that Nvidia's CUDA moat extends fully to the inference layer.
NoveltySignificance
AI Infra · Player MapSilicon Substrate
Moonshot AI
Moonshot AI

Foundation Models / LLMs

View Company Profile

Moonshot AI's Kimi K3 hits 709 tokens/sec on 16 Google TPU v7 chips, beating GB200's 452 in a like-for-like inference test built by vLLM's founding team.

Inferact, a startup founded by the original vLLM maintainers, benchmarked 16 Google TPU v7 Ironwood chips against 16 Nvidia GB200 GPUs running the same Kimi K3 model on the same vLLM inference engine. The TPU configuration reached 709 tokens per second versus 452 on GB200 — a 57% gap — attributed entirely to a hand-written "megakernel" that collapses hundreds of individually scheduled kernels into a single Pallas program. On Qwen 3.8 27B, four TPU v7 chips hit 1,515 tokens/sec versus 695 on equivalent GB200 hardware. With speculative decoding disabled, TPU still nearly doubled GB200 at batch size 1 (249 vs 127 tokens/sec). Accuracy was unchanged: Kimi K3 scored 94.4% on GPQA-Diamond and 97.2% on GSM8K on both platforms. The speed advantage is notable because GB200 actually has higher HBM bandwidth (8,000 GB/s vs TPU v7's 7,380 GB/s), so the gap stems from software orchestration rather than raw memory throughput. Inferact's megakernel leverages TPU v7's 64 MiB per-TensorCore VMEM, which is software-managed, enabling cross-layer weight prefetching that Nvidia's hardware-scheduled 38 MiB distributed SM cache makes harder. The code is open-sourced as tpu-megakernels under a joint engineering partnership with Google Cloud.

The result matters because it attacks Nvidia's inference moat from an unexpected angle: not with cheaper or more powerful silicon, but with a software layer written by engineers who built their reputations on CUDA-based systems. TPU v7 has long been viewed as competitive on training cost but behind on inference flexibility; megakernels shift that calculus. For Moonshot AI, the benchmark is a distribution signal as much as a performance one — Kimi K3 already went live on Amazon Bedrock in late September, and demonstrating first-class performance across both Nvidia and Google silicon makes the model more portable for enterprise buyers wary of single-vendor lock-in. The involvement of DSpark, DeepSeek's speculative-decoding framework, also underscores how Chinese open-weight research is now circulating through Western inference stacks. Inferact's $150M seed round at an $800M valuation, led by a16z and Lightspeed and closed earlier this year, looks prescient if TPU inference optimization becomes a standalone category.

The practical implication for builders is that inference cost per token is becoming a software problem, not just a hardware procurement decision. Teams deploying large MoE models should watch whether megakernel-style approaches generalize beyond Kimi K3's specific architecture — Inferact says they plan to extend support to more model structures, but the current kernel is custom-tuned. For investors, the more interesting question is whether Google Cloud can convert TPU inference wins into enterprise workload migration, or whether Nvidia's CUDA ecosystem and broader model compatibility keep buyers anchored. The open-source release lowers the barrier for anyone running TPU workloads to test the approach directly.

#MoonshotAI #KimiK3 #GoogleTPU #NvidiaGB200 #AIInference #vLLM

#Moonshot AI#Kimi K3#Google TPU v7#Nvidia GB200#AI inference optimization#related:Google#related:Nvidia#related:DeepSeek

How This Connects

Based on AI Infra · Player Map

  1. 1d agoMoonshot AI's Kimi K3 hits 709 tokens/sec on 16 Google TPU v7 chips, beating GB200's 452 in a like-f... · THIS ARTICLE
  2. 2d agoDensityAI, the AI chip startup founded a year ago by former Tesla Dojo leaders, is in late-stage tal...DensityAI
  3. 4d agoCoreWeave completes $4.2 billion convertible note offering for AI data-center expansionCoreWeave
  4. 6d agoNvidia works to ease the electrical power bottleneck slowing AI data center expansionNvidia's
  5. 2w agoPositron AI closes $875M Series C at $5B to tape out LPDDR5X inference ASIC AsimovPositron AI
  6. 3w agoNVIDIA to Invest $3.5 Billion in MediaTek in Expanded AI Chip PartnershipNVIDIA invests $3.5B in MediaTek

Related News

More news from Moonshot AI

Stay updated with the latest news and announcements from Moonshot AI.

View all Moonshot AI news

Discover AI Startups

Explore 5,000+ AI companies with VC-grade analysis, funding data, and investment insights.

Explore Dashboard