Beta The Briev beta is out. Free on iPhone via TestFlight — install it in under a minute.

Join the beta ↗
Briev
Live
Technology

Nvidia and Cerebras tout record inference speeds that customers may never use

Nvidia announced its Groq-3-based LPX racks can process 3,400 tokens per second, a figure four times faster than Cerebras, which responded with comparable claims for its CS-4 accelerators. Analysts say these headline numbers reflect peak performance that most users will not achieve in real deployments.

During the Hot Chips event in California, Nvidia unveiled its Groq-3-based LPX racks, claiming they can generate 3,400 tokens per second on the Gemma 4 31B model, a rate Nvidia says outpaces Cerebras by a factor of four. Cerebras quickly replied with performance figures for its next-generation CS-4 accelerators that appear comparable, though both sets of numbers stem from single-request benchmarks performed by Artificial Analysis.

Analysts note that while the raw speed is impressive, the SRAM-centric architectures lack sufficient memory to handle larger batches, limiting a single rack to roughly twelve concurrent 100,000-token inputs. This constraint means the headline speeds are unlikely to be realized in typical inference-as-a-service scenarios, where cost, power use, and scalability matter more. The article argues that combining these accelerators with GPUs—an approach Nvidia, AMD, and AWS are pursuing—could mitigate the memory bottleneck and make “premium inference” economically viable. Until such hybrid benchmarks are released, the touted figures remain more marketing than practical performance indicators.

Why it matters

Understanding the real-world limits of AI inference hardware helps businesses gauge true cost and performance expectations.

In this story

Groq-3LPX rackCS-4 acceleratortoken throughputinference latencySRAM architecturebatch sizehybrid GPU accelerationAI hardware benchmark
Get the beta ↗