
OpenAI's Jalapeno Chip Claims Faster Inference Than Nvidia Systems
The AMW Read
OpenAI's first reported custom inference-chip benchmarks materially extend its case study and could reshape serving economics beyond a single model product.
OpenAI's Jalapeno Chip Claims Faster Inference Than Nvidia Systems
OpenAI says its Jalapeno inference ASIC, developed with Broadcom, delivered lower latency and more work per watt than comparison systems based on Nvidia's GB200 or GB300 superchips in InferenceX tests. Across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, OpenAI reported 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower end-to-end latency. The company plans small-volume deployment by year-end and a larger ramp in 2027, while retaining Nvidia as part of its broader compute strategy.
The result matters less as a declaration of a wholesale Nvidia replacement than as evidence that frontier labs are extending vertical integration into inference hardware. Faster token delivery and higher efficiency address the two operating constraints behind interactive models and agents: responsiveness for users and energy cost at scale. But the figures remain OpenAI's benchmark claims, and the planned initial deployment is limited; a fleet-level advantage will depend on reproducibility across workloads, operational reliability, and manufacturing scale.
For builders, the practical signal is to evaluate serving partners on latency and work-per-watt under their own model mix rather than peak throughput alone. For investors, Jalapeno raises the strategic value of custom silicon programs that can tune inference economics for a lab's workloads, while also showing why Nvidia's installed ecosystem remains difficult to displace quickly.

