Skip to main content
Back to News
NVIDIA Groq 3 LPX enters full production for ultrafast agentic inference
Technology
2 min read
US

NVIDIA Groq 3 LPX enters full production for ultrafast agentic inference

The AMW Read

Full production of Groq-derived LPX plus Nebius Token Factory adoption meaningfully advances known NVIDIA–Groq silicon integration and is structural for inference economics and neo-cloud differentiation.
NoveltySignificance
AI Infra · Player MapSilicon Substrate

Named counterparties: Groq, Nebius Group

NVIDIA Groq 3 LPX enters full production for ultrafast agentic inference

NVIDIA announced at Hot Chips that Groq 3 LPX, its interactive AI inference accelerator extending the Vera Rubin platform, is now in full production. The rack-scale system targets ultrafast token generation for agentic workloads. In Artificial Analysis benchmarks, Groq 3 LPX hit 3,400 output tokens per second on the open-source Gemma 4 31B model at a 100,000-token context — the fastest recorded for that model — and NVIDIA claims roughly 4x faster responsiveness versus the nearest alternative for agents and latency-sensitive jobs. Nebius is the first AI cloud to adopt the hardware, planning to expose it through Nebius Token Factory; Groq's own inference cloud says it will also be among the earliest adopters.

Why it matters sits at the intersection of AI Infrastructure (Segment 04) and the silicon substrate: NVIDIA is productizing the Groq LPU lineage as a Vera Rubin companion rather than a GPU replacement, explicitly splitting processing enormous context from generating tokens at extreme interactivity. That codesign move updates the neo-cloud player map — Nebius differentiating Token Factory on generation latency — and feeds the Segment 04 debate over whether specialized inference silicon compounds Nvidia allocations or gets absorbed into the NVL stack. It also tightens the loop for agentic systems, where hundreds or thousands of inference steps make tokens-per-second the binding constraint on real-time tool use and coding agents.

For builders and investors, the near-term signal is commercial availability of LPX-class interactivity through Nebius APIs without a stack migration, while longer-term the question is whether dedicated generation accelerators become a standard AI-factory SKU beside Vera Rubin NVL72 — raising the bar for pure-GPU neo-clouds and for challenger silicon that remains outside Nvidia's licensed platform.

#NVIDIA #Groq3LPX #Nebius #AIInfrastructure #Inference #AgenticAI

#NVIDIA#Groq 3 LPX#Nebius#Vera Rubin#inference accelerator#agentic AI#related:Nebius#related:Groq
Read Original

How This Connects

Based on AI Infra · Player Map

  1. 1d agoNVIDIA Groq 3 LPX enters full production for ultrafast agentic inference · THIS ARTICLE
  2. 1w agoEtched on Tuesday announced a $700 million raise at a $21 billion valuation, led by Jane Street afte...Etched
  3. 2w agoIntel raised $20 billion in an upsized share offering priced at $95 per share, a 2.6% discount to Mo...Intel
  4. 2w agoSource Foundry, a semiconductor startup founded by former engineers, has raised an additional $400 m...Source Foundry

Related News

Discover AI Startups

Explore 5,000+ AI companies with VC-grade analysis, funding data, and investment insights.

Explore Dashboard