Huawei is accelerating the development of its Ascend AI accelerator family as it tries to build a larger alternative to Nvidia’s computing ecosystem in China. At Huawei Connect 2026, the company outlined an annual chip cadence, introduced new SuperPoD systems and moved the planned release of two Ascend 960 variants forward.
Huawei says Ascend 960DT, aimed at AI training, will be available in the first quarter of 2027, while Ascend 960PR, aimed at inference, is planned for the third quarter. The company had previously placed the 960DT later on its roadmap. Huawei also said future Ascend 970 and Ascend 980 chips are planned for 2028 and 2029.
The strategy is broader than producing a faster individual accelerator. Huawei is attempting to solve the same system-level problems that have become central to large AI deployments: memory capacity, interconnect bandwidth, power consumption, software support and the ability to combine large numbers of processors into one computing system.
Scaling matters as much as the chip
Huawei’s new Ascend 960 SuperPoD is designed to scale to 4,096 accelerators. The company says the system uses its UnifiedBus architecture and Hi-ONE near-packaged optical interconnect technology. Huawei claims the optical system can provide 7.2 Tbit/s transmission capacity per engine.
According to Huawei, an Ascend 960 SuperPoD can reach 8 exaflops of FP8 compute performance and 16 exaflops of FP4 performance. The company also says its design can reduce power consumption by more than 550 kilowatts compared with a configuration requiring tens of thousands of conventional 800G optical modules.
These are vendor-provided figures, and comparisons with Nvidia systems depend heavily on workload, precision, software and system configuration. Raw accelerator performance is only one part of an AI cluster. Memory bandwidth and capacity, networking, compiler support and the ability to keep thousands of processors busy can determine how much of that theoretical performance becomes useful work.
China’s demand is becoming a constraint
Reuters has reported that demand for Huawei’s AI computing equipment in China is exceeding available production, limiting the company’s ability to expand internationally. That creates a different problem from simply being unable to compete on specifications: Huawei has to build enough systems to satisfy customers while continuing to improve the hardware.
The supply issue also reflects the wider semiconductor constraints surrounding China’s AI industry. Restrictions on advanced AI processors and high-bandwidth memory have encouraged Chinese companies to develop domestic alternatives, but domestic production still has to overcome manufacturing and component limitations.
Huawei’s response is to control more of the stack. The company is developing accelerators, networking, optical interconnects, storage and system software. It has also been opening parts of its software ecosystem to developers and says Ascend supports more than 90 leading open-source projects.
The software battle remains difficult
Nvidia’s advantage is not limited to its GPUs. CUDA and the surrounding software ecosystem have become deeply embedded in AI development. Developers have spent years building models, kernels, libraries and deployment systems around Nvidia hardware.
Huawei therefore needs developers to see Ascend as more than a government-supported substitute. The company has been working on CANN, PyTorch integration and other tools designed to make model migration easier. It has also opened large shared compute resources to developers and announced additional investment in the Ascend ecosystem.
The competitive picture is still developing. Huawei does not need to reproduce Nvidia’s entire global business to matter in China. It needs a sufficiently capable combination of chips, interconnects, software and supply to support Chinese AI companies at scale.
The accelerated Ascend road map shows that Huawei is treating AI computing as a continuing infrastructure race. The next important evidence will come from actual deployments, production volumes, developer adoption and independent measurements of how the systems perform on real training and inference workloads.