
NVIDIA Groq 3 LPX enters full production for ultrafast agentic inference
The AMW Read
Full production of Groq-derived LPX plus Nebius Token Factory adoption meaningfully advances known NVIDIA–Groq silicon integration and is structural for inference economics and neo-cloud differentiation.
Named counterparties: Groq, Nebius Group
NVIDIA Groq 3 LPX enters full production for ultrafast agentic inference
NVIDIA announced at Hot Chips that Groq 3 LPX, its interactive AI inference accelerator extending the Vera Rubin platform, is now in full production. The rack-scale system targets ultrafast token generation for agentic workloads. In Artificial Analysis benchmarks, Groq 3 LPX hit 3,400 output tokens per second on the open-source Gemma 4 31B model at a 100,000-token context — the fastest recorded for that model — and NVIDIA claims roughly 4x faster responsiveness versus the nearest alternative for agents and latency-sensitive jobs. Nebius is the first AI cloud to adopt the hardware, planning to expose it through Nebius Token Factory; Groq's own inference cloud says it will also be among the earliest adopters.
Why it matters sits at the intersection of AI Infrastructure (Segment 04) and the silicon substrate: NVIDIA is productizing the Groq LPU lineage as a Vera Rubin companion rather than a GPU replacement, explicitly splitting processing enormous context from generating tokens at extreme interactivity. That codesign move updates the neo-cloud player map — Nebius differentiating Token Factory on generation latency — and feeds the Segment 04 debate over whether specialized inference silicon compounds Nvidia allocations or gets absorbed into the NVL stack. It also tightens the loop for agentic systems, where hundreds or thousands of inference steps make tokens-per-second the binding constraint on real-time tool use and coding agents.
For builders and investors, the near-term signal is commercial availability of LPX-class interactivity through Nebius APIs without a stack migration, while longer-term the question is whether dedicated generation accelerators become a standard AI-factory SKU beside Vera Rubin NVL72 — raising the bar for pure-GPU neo-clouds and for challenger silicon that remains outside Nvidia's licensed platform.



