INFRA Signal 442
Cerebras CS-4 racks pair WSE-3T 'Turbo' wafers with AWS Trainium XPUs and AMD Instinct GPUs
Cerebras CS-4 rack systems pair WSE-3T wafers, same 5nm silicon pushed to roughly double the clock, with AWS Trainium XPUs and AMD Instinct GPUs so Cerebras handles decode while partners handle prompt processing.
The WSE-3T is not new silicon; it is the existing WSE-3 with improved power delivery enabling higher clocks at roughly twice the TDP, doubling compute, memory fabric and I/O bandwidth. Cerebras targets a 10x throughput-per-watt improvement over the previous generation, but the published numbers rely on sparse FP16 and dense LLM inference is unlikely to saturate the on-chip SRAM. The CS-4 signals a strategic retreat from running the full inference stack on its own silicon in favor of a disaggregated pipeline that mirrors how Nvidia deploys Groq LPUs.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
The WSE-3T reuses the same TSMC 5nm wafer as the WSE-3 to 4 trillion transistors, 900,000 cores, 44 GB SRAM, but doubles sparse FP16 to 250 PFLOPS and memory bandwidth to 43.2 PB/s by raising the clock to roughly 2.8 GHz and doubling TDP to 33 kW per wafer and 46 kW per system.
CS-4 stops trying to run the full inference stack on Cerebras silicon; AWS Trainium XPUs and AMD Instinct GPUs handle prompt processing while WSE-3T wafers serve as decode accelerators, mirroring how Nvidia integrates Groq LPUs into LPX racks.
Once sparsity is removed the dense FP16 figure is closer to 25 PFLOPS, and the prior-generation WSE-3 was already unable to saturate its SRAM bandwidth on real LLM inference workloads, a limitation the Turbo refresh does not address.
THE CLUSTER