Huawei's next-gen Ascend NPUs could become China's best option
At its annual Connect conference, Huawei unveiled a new generation of AI accelerators that promise performance far beyond anything Nvidia can currently sell in the Middle Kingdom.
This could be a boon for Chinese developers used to purchasing lesser weapons from everyone's favorite AI arms dealer.
If that weren’t enough, the chip, codenamed the Ascend 960DT — the DT here apparently stands for decode/training — is slated to arrive a full three quarters ahead of schedule, launching in the first quarter of 2027.
With up to 288 GB of what we assume is Huawei’s custom HiZQ memory tech, a homegrown alternative to the high-bandwidth memory used by the rest of the world, and up to four petaFLOPS of FP4 performance (half that at FP8), the accelerator offers twice the performance and memory capacity of the company’s 950-series parts, launched earlier this year.
Compared to American GPUs, the 960DT offers similar memory and bandwidth to Nvidia’s B300 family of chips launched last year, but only about half the FP8 and a third the FP4 compute.
While a big step up for the Chinese chip designer, it still has a long way to go to catch up with Nvidia’s Rubin and AMD's recently launched MI455X which promise substantially higher compute and bandwidth.
Rubin boasts between 35 and 50 petaFLOPS of FP4 performance, 288 GB of HBM4 memory, and 22 TB/s of bandwidth, putting it in an entirely different league.
Having said that, neither of those chips is available for sale in the Middle Kingdom, so it's not like Chinese model devs have better options.
The best GPU Nvidia can sell in China offers near-identical dense floating-point performance, twice the memory, double the memory bandwidth, and support for much larger scale-up domains.
Scaling up The performance gap may not be nearly as damning as it sounds, either.
AI models aren’t trained, and for the most part aren’t run on a single GPU or NPU anymore.
The more important factor is often how efficiently the platform scales.
Nvidia's and AMD's latest systems pack 72 GPUs into a single rack-scale compute platform, which can be expanded to 576 using optical interconnects for scale up and scale out networking.
For its upcoming 960-series accelerators, Huawei is going a similar route using near packaged optics (NPO) to scale its compute domain to as many as 4,096 chips capable of delivering up to 16 exaFLOPS of FP4 compute.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.