Huawei is not trying to catch Nvidia one chip at a time. Its latest strategy is to connect thousands of processors into increasingly large systems and compete at the infrastructure level instead.
The AI giant unveiled the Atlas 960E SuperPoD and accelerated the launch of its Ascend 960DT AI chip to the first quarter of 2027. The system can scale to 4,096 NPUs and is designed to integrate computing, networking, memory, storage, and software into a single AI infrastructure platform.
Huawei is positioning the stack as an alternative to Nvidia, but limited chip supply and Nvidia’s widely used CUDA software remain major barriers to broader adoption.
Huawei accelerates its Ascend 960 roadmap
According to Huawei, development of Ascend 960 has exceeded expectations. The Ascend 960DT now scheduled for the first quarter of 2027, three quarters earlier than originally planned. Meanwhile, the Ascend 960PR is expected in the third quarter.
The company also plans to introduce a new Ascend generation each year, with the Ascend 970 and 980 scheduled for 2028 and 2029, respectively.
Reuters reported that Huawei is already struggling to produce enough AI computing equipment to satisfy demand in China. Rotating chairman Eric Xu said the company therefore has no plans for a large-scale overseas expansion for now.
Atlas 960E puts scale and interconnects at the center
The Atlas 960E SuperPoD combines Ascend 960 chips with Huawei’s UnifiedBus interconnect and Hi-ONE near-packaged optics technology.
According to Huawei, a single system can scale to 4,096 NPUs and deliver 8 EFLOPS of FP8 compute performance with up to 1 petabyte of high-bandwidth memory. Huawei said the system uses 5,500 Hi-ONE optical engines instead of the 48,000 800G optical modules that would traditionally be needed to connect the NPUs.
The company claims the design cuts power consumption by more than 550 kilowatts and provides 99.8% system availability.
The 4,096-NPU figure applies to a single SuperPoD, not Huawei’s larger scaling ambitions. Huawei said multiple SuperPoDs can be connected into a SuperCluster supporting up to 512,000 NPUs, while a multi-rail topology could eventually extend that architecture to 1 million NPUs.
TechCrunch noted that Huawei had previously discussed an Atlas 960 system containing 15,488 Ascend 960 chips. The different configurations underscore Huawei’s broader emphasis on scaling AI compute through increasingly large interconnected systems.
For now, Huawei’s strategy is less about relying on a single accelerator and more about improving aggregate performance by connecting large numbers of processors.
Huawei is also building out its software ecosystem
Nvidia’s biggest advantage is not limited to its GPUs. Its CUDA software platform has become deeply embedded in AI development, creating another hurdle for competing hardware providers.
Huawei is trying to narrow that software gap through CANN, its Compute Architecture for Neural Networks.
The company said CANN has moved to sustained, community-driven open-source development and now has more than 5,200 monthly active developers. Huawei also highlighted that Ascend supports more than 90 third-party open-source projects, including PyTorch, Triton, vLLM, and veRL. More than 40 models have been natively pretrained on Ascend and CANN.
Reuters noted that Nvidia still retains a major software advantage through CUDA, despite Huawei’s efforts to grow Ascend adoption.
Huawei’s infrastructure push creates a bigger channel opportunity
Huawei’s strategy extends the AI hardware race well beyond accelerator chips. Systems containing thousands of NPUs also require high-speed optical networking, storage, power management, cooling, and increasingly sophisticated integration between those components.
That could create a broader opportunity for infrastructure vendors and partners as AI deployments become larger and more complex. Channel Insider has previously reported that AI adoption is exposing network readiness gaps, creating demand for partners capable of modernizing the infrastructure supporting increasingly data-intensive workloads.
Power and cooling are becoming part of that equation as well. As AI clusters become denser, data center operators are facing higher energy and thermal requirements, with AI-ready cooling becoming another part of the infrastructure stack that infrastructure partners may need to address.
Customers evaluating systems at this scale are therefore not simply choosing processors. They are weighing an entire stack, including networking capacity, software compatibility, energy requirements, support, and suppliers’ ability to deliver hardware at the required scale.
Huawei still faces substantial hurdles. Nvidia’s CUDA ecosystem remains well-established, while Huawei has acknowledged that its current production capacity is insufficient to meet demand in China, limiting its plans for large-scale overseas expansion.
The Atlas 960E therefore matters as much for its architecture as its raw specifications. If AI competition continues to shift toward massive interconnected systems, the companies supplying networking, optics, storage, cooling, and integration for those accelerators could become increasingly important to how the AI infrastructure market develops.
DeepSeek reportedly plans to deploy 160,000 Huawei AI chips at a data center in Inner Mongolia, a project that could further expand Huawei’s role in China’s AI infrastructure market.





