Wednesday, 19 August 2026 SourcesAbout🌓
🇬🇧 UK ▾
BREAKING
Technology

Cerebras CS-4 rack systems juice chips for every last drop of AI performance

The Register ·
Cerebras CS-4 rack systems juice chips for every last drop of AI performance

If high-speed AI inference is what you’re after, memory bandwidth is the bottleneck to beat.

At a mind-numbing 21.6 petabytes per second (PB/s) of memory bandwidth, Cerebras' dinner-plate-sized AI accelerators were already 1,000x faster than Nvidia's or AMD’s best GPUs.

The chip newcomer unveiled its next-gen Wafer Scale Engine (WSE) and Nexus rack systems on Tuesday.

Cerebras aims to extend that lead by boosting throughput per watt tenfold over the previous generation.

Putting the 'T' in Turbo Cerebras accomplishes this in a couple of ways.

But, from what we can tell, the primary lever comes from squeezing its chips for every hertz they’ve got.

The newly announced WSE-3T — the “T” here stands for “Turbo” — promises twice the compute, memory fabric, and I/O bandwidth of the now two-year-old WSE-3.

Yet, if you look at the chart below, you’ll notice it accomplishes this using the same process tech, wafer area size, transistor count, core count, and SRAM capacity.

That's because the WSE-3T isn't new silicon.

Instead, Cerebras tells us it's just pushing its existing wafer scale engine harder.

The main innovation this time around seems to be related to power delivery, which is apparently so efficient that they’re able to push twice the power through the chip, which “enables higher operating frequencies and faster token generation.” How much higher does it clock? By our estimate, Cerebras is now running the silicon at 2.8 GHz, up from 1.4 GHz last gen, which would be quite the accomplishment.

In any case, each WSE-3T boasts 250 petaFLOPS of AI compute, 44 GB of SRAM (that’s not a typo, there really is that much SRAM on there), good for 43.2 PB/s of memory bandwidth, and 2.4 Tbps of off-die connectivity.

On paper that sounds more impressive than it really is.

AMD and Nvidia’s latest GPUs offer 4 to 5 petaFLOPS of dense FP16 compute or 35 to 50 petaFLOPS at FP4.

Read the full article on The Register ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.

More from The Register

See all ›

More in Technology

See all ›