Huawei launched the Ascend 960 SuperPoD at its 2026 Connect conference, unveiling an ultra-large-scale AI computing cluster designed to function as a single logical machine. The system represents Huawei’s aggressive push toward independent computing infrastructure free from US semiconductor restrictions. The SuperPoD combines 15,488 Ascend AI processing units across 220 cabinets occupying 2,200 square meters of physical space.
The Ascend 960 represents the culmination of Huawei’s three-year AI chip roadmap spanning 2026 through 2028. The roadmap includes the Ascend 950 series launching in 2026, the Ascend 960 in Q4 2027, and the Ascend 970 in 2028, with each generation doubling computing capacity. Eric Xu, Huawei’s rotating chair, announced that the Atlas 950 SuperPoD would deliver 6.7 times greater computing power and 15 times more memory capacity than NVIDIA’s corresponding generation, reaching 1,152 terabytes of total memory.

Two technological innovations distinguish the Ascend 960 from previous generations. Near-Packaged Optics (NPO) represents an integrated light source enabling faster optical connections while reducing energy consumption. The upgraded UnifiedBus interconnect technology allows simultaneous connection of 4,000 processors into a unified computing framework.
The Ascend 950 series features new low-precision formats, stronger vector processing with refined memory access granularity from 512 bytes to 128 bytes, and a 2 terabyte-per-second interconnect bandwidth: 2.5 times that of the Ascend 910C. Together, these technologies enable the system to achieve three to four times higher training and inference performance compared to earlier SuperPoD generations.
David Wang Tao, Huawei’s rotating and acting chairman, emphasized that supernodes represent inevitable progression toward ultra-large-scale data centers. Huawei announced plans for mass production of the industry’s first NPO product with integrated light sources, signaling commercial viability rather than laboratory experimentation.

The Ascend 950PR, available in card and SuperPoD server formats starting Q1 2026, supports prefill, recommendation, and agent-based applications with HiBL 1.0, Huawei’s proprietary low-cost HBM technology more cost-effective than competing solutions like HBM3E and HBM4E.
The company simultaneously filed research on arXiv describing a 256K-node-class single AI computing framework extending capabilities far beyond individual SuperPoD installations.
The architectural approach differs fundamentally from distributed systems. Huawei is developing a large-scale AI computing cluster called SuperPoD, similar to NVIDIA’s DGX SuperPOD but using Huawei’s proprietary Ascend series AI chips, and has introduced SuperCluster architecture that connects multiple SuperPoD units. Huawei describes the expanded framework as operating as one unified machine rather than a network of over 200,000 separate machines.

This design addresses latency and synchronization challenges that plague conventional distributed computing when scaling to extreme processor counts. The previously announced Atlas 900 A3 SuperPoD, equipped with 384 Ascend 910C NPUs, achieved processing power of 300 PFLOPS, demonstrating Huawei’s escalating capability trajectory. However, building such systems without access to NVIDIA GPUs or advanced US semiconductor technology requires alternative component strategies.
The Ascend 960 builds directly upon the Atlas 950, which houses 8,192 Ascend NPUs spanning 128 computing racks across approximately 1,000 square meters. The Ascend 950DT, set for launch in Q4 2026, targets training and decoding workloads with higher demands for memory and bandwidth, integrating HiZQ 2.0 HBM offering 144 gigabytes of memory and 4 teabyte-per-second memory access bandwidth.
The generation jump from the 950 to 15,488 Ascend 960 chips represents nearly a doubling of computing density within the same architectural framework. Huawei indicated that Ascend chip releases will continue following an aggressive upgrade cycle guided by the Tau Scaling Law, positioning each new iteration to push performance beyond the previous generation’s capabilities.


















