Cerebras unveils turbo-boosted WSE-3T chips and modular CS-4 rack system
Cerebras introduced the WSE-3T accelerator, a faster version of its wafer-scale engine, and the CS-4 rack that houses up to three of these units.
Cerebras announced its next-generation Wafer Scale Engine, the WSE-3T, which it brands as “Turbo” because it pushes the same 46,225 mm², 4-trillion-transistor die harder, raising clock speed to about 2.8 GHz. This results in twice the compute capacity, memory bandwidth (43.2 PB/s) and I/O throughput (2.4 Tbps) of the prior WSE-3, while power consumption climbs to an estimated 33 kW per wafer and up to 46 kW per system.
The accelerator’s 44 GB of SRAM enables it to act mainly as a decode engine, with pre-fill work offloaded to AWS Trainium and AMD Instinct GPUs. Alongside the chip, Cerebras unveiled the CS-4 rack, a modular chassis that holds up to three “backpack” units, each containing an accelerator, and separates power shelves for easier scaling. The rack’s total power draw is projected between 120 kW and 140 kW, still lower than upcoming 240-250 kW systems from competitors. Inter-chip latency has been cut from five to two microseconds by eliminating switches and using a 2-D torus mesh, supporting models up to 50 trillion parameters and delivering up to 4,400 tokens per second on a single CS-4 system.
Why it matters
The launch shows how AI hardware is evolving toward higher efficiency and modularity, influencing future data-center designs.
In this story
