Friday, 28 August 2026 SourcesAbout🌓
🇬🇧 UK ▾
BREAKING
Technology

Nvidia and Cerebras are selling performance their customers will (probably) never see

The Register ·
Nvidia and Cerebras are selling performance their customers will (probably) never see

Shots were fired at the Hot Chips conference in California this week as Nvidia announced its new Groq-3-based LPX racks had entered production, with early tests showing the systems churning an eye-watering 3,400 tokens a second in Gemma 4 31B.

That's four times faster than rival Cerebras.

A day later, Cerebras fired back, touting nearly equivalent performance from its next-gen CS-4 accelerators revealed last week.

These top-line performance figures make inference feel instantaneous relative to the chatbots we've grown accustomed to over the past four years.

But the two companies are essentially arguing over numbers their customers will probably never see in production.

To be clear, neither is lying.

If you wanted to recreate these results, you certainly could — Artificial Analysis, the team responsible for both sets of benchmarks, knows what they're doing — but beyond a marketing gimmick, no inference-as-a-service model operator in their right mind would run either system this way.

Not unless you've somehow figured out how to make a profit serving one request at a time.

In reality, these numbers are more like the top speed on a race car.

It makes for great marketing but you probably aren't driving that fast on a regular basis, and if you did, you wouldn't get very far before your tank runs dry.

These aren't the numbers that matter Unsurprisingly, the economics of "premium" or "ultra-low latency" inference are a bit more nuanced, but ultimately come down to how efficiently you can scale that performance at the rightmost end of the Pareto frontier.

We've shown the graphic below before.

It charts the performance characteristics of how different Nvidia B300 configurations perform across a Pareto front.

As you can see, GPUs are great for high-throughput, low-interactivity applications, but run out of steam quite quickly as per-user generation rates climb.

Read the full article on The Register ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.

More from The Register

See all ›

More in Technology

See all ›