Moonshot AI's Kimi K3 hits 709 tokens/sec on 16 Google TPU v7 chips, beating GB200's 452 in a like-f...
The AMW Read
First credible benchmark showing TPU inference beating Nvidia on identical models and engines via software, challenging the assumption that Nvidia's CUDA moat extends fully to the inference layer.
Moonshot AI's Kimi K3 hits 709 tokens/sec on 16 Google TPU v7 chips, beating GB200's 452 in a like-for-like inference test built by vLLM's founding team.
Inferact, a startup founded by the original vLLM maintainers, benchmarked 16 Google TPU v7 Ironwood chips against 16 Nvidia GB200 GPUs running the same Kimi K3 model on the same vLLM inference engine. The TPU configuration reached 709 tokens per second versus 452 on GB200 — a 57% gap — attributed entirely to a hand-written "megakernel" that collapses hundreds of individually scheduled kernels into a single Pallas program. On Qwen 3.8 27B, four TPU v7 chips hit 1,515 tokens/sec versus 695 on equivalent GB200 hardware. With speculative decoding disabled, TPU still nearly doubled GB200 at batch size 1 (249 vs 127 tokens/sec). Accuracy was unchanged: Kimi K3 scored 94.4% on GPQA-Diamond and 97.2% on GSM8K on both platforms. The speed advantage is notable because GB200 actually has higher HBM bandwidth (8,000 GB/s vs TPU v7's 7,380 GB/s), so the gap stems from software orchestration rather than raw memory throughput. Inferact's megakernel leverages TPU v7's 64 MiB per-TensorCore VMEM, which is software-managed, enabling cross-layer weight prefetching that Nvidia's hardware-scheduled 38 MiB distributed SM cache makes harder. The code is open-sourced as tpu-megakernels under a joint engineering partnership with Google Cloud.
The result matters because it attacks Nvidia's inference moat from an unexpected angle: not with cheaper or more powerful silicon, but with a software layer written by engineers who built their reputations on CUDA-based systems. TPU v7 has long been viewed as competitive on training cost but behind on inference flexibility; megakernels shift that calculus. For Moonshot AI, the benchmark is a distribution signal as much as a performance one — Kimi K3 already went live on Amazon Bedrock in late September, and demonstrating first-class performance across both Nvidia and Google silicon makes the model more portable for enterprise buyers wary of single-vendor lock-in. The involvement of DSpark, DeepSeek's speculative-decoding framework, also underscores how Chinese open-weight research is now circulating through Western inference stacks. Inferact's $150M seed round at an $800M valuation, led by a16z and Lightspeed and closed earlier this year, looks prescient if TPU inference optimization becomes a standalone category.
The practical implication for builders is that inference cost per token is becoming a software problem, not just a hardware procurement decision. Teams deploying large MoE models should watch whether megakernel-style approaches generalize beyond Kimi K3's specific architecture — Inferact says they plan to extend support to more model structures, but the current kernel is custom-tuned. For investors, the more interesting question is whether Google Cloud can convert TPU inference wins into enterprise workload migration, or whether Nvidia's CUDA ecosystem and broader model compatibility keep buyers anchored. The open-source release lowers the barrier for anyone running TPU workloads to test the approach directly.

