Nvidia's Groq-Powered LPX Racks Claim Four-Fold Speed Edge Over Competitors
Nvidia reported that its Groq 3-based LPX rack achieved 3,400 tokens per second on the Gemma 4 31B model, roughly four times faster than the closest rival platform.
Nvidia unveiled its first performance data for the Groq 3-based LPX rack, announcing a throughput of 3,400 tokens per second on Google’s Gemma 4 31B model with a 100,000-token input sequence, according to an Artificial Analysis test. The company claims this represents a four-fold advantage over the nearest alternative, Cerebras, which managed 882 tok/s in the same benchmark. Groq’s third-generation LPUs employ an SRAM-centric dataflow design, offering roughly 150 TB/s of memory bandwidth but only 500 MB of on-die memory per chip, requiring models to be distributed across up to 256 LPUs per rack and multiple racks to be linked via Ethernet.
Nvidia’s strategy pairs these LPUs with its Vera Rubin GPUs, assigning the compute-intensive prefill phase to GPUs and the memory-bandwidth-heavy decode phase to LPUs, aiming to improve interactivity for AI agents. Early adopter Nebius in the Netherlands will be among the first to field the combined system. Analysts caution that while the benchmark is impressive, it reflects a best-case scenario for a relatively small dense model, and scaling to larger MoE models could demand many more accelerators, potentially narrowing the gap with rivals such as Cerebras’s upcoming CS-4 chips.
Why it matters
The results highlight Nvidia's push for faster AI inference hardware, which could set new standards for responsive AI services.
In this story
