DeepSeek is expanding its work with Huawei, releasing open-source software designed for the Chinese tech company’s Ascend processors.
DeepSeek’s Sept. 30 releases and updates extend Ascend support across six open-source projects covering low-level operations needed to build and optimize AI workloads. The release moves their collaboration further into the software layer around AI infrastructure.
For channel partners, the practical question is how much existing development work can carry over when customers evaluate another accelerator platform.
Shared interfaces reduce some redevelopment
Six components now support Huawei’s chip platform.
Two releases are especially relevant to developers moving between hardware platforms. DeepGEMM-Ascend retains the API and development workflow used by DeepGEMM on other hardware, while TileKernels exposes the same Python APIs across Nvidia GPUs and Huawei NPUs.
Four more components extend support across different parts of the stack:
- TileLang adds native support for the Ascend 950, including tools for writing and optimizing low-level AI operations.
- DeepEP-Ascend handles communication between accelerators, including data transfers needed for large mixture-of-experts models.
- FlashMLA adds sparse-attention kernels for processing prompts and generating tokens on Ascend 950 NPUs.
- DeepSelect adds TopK operations for the platform to select relevant data during model processing.
DeepSeek’s repositories credit Huawei with technical and engineering support. Software maturity has been a key hurdle for the chip platform as it competes with an Nvidia ecosystem built around years of CUDA development.
Huawei also says it has upgraded Ascend C for the Ascend 950 generation and made progress opening its PTO instruction set through a wider developer tooling effort.
Collaboration reaches a 128-chip Ascend system
DeepSeek said it worked with Huawei to optimize computation and communication for a supernode configuration based on 128 Ascend 950 chips, according to Reuters.
DeepSeek reports that DeepEP-Ascend reached roughly 90% to 95% of the physical payload bandwidth limit in expert-parallel dispatch tests with up to 32 participating ranks. Larger configurations and combine operations remain under optimization.
The company has also been linked to a planned 160,000-chip Ascend deployment in Inner Mongolia, extending the relationship from software and cluster optimization into plans for much larger AI infrastructure.
Partners still have validation work to do
For systems integrators, shared APIs can reduce some redevelopment without making Nvidia and Ascend environments interchangeable. Teams evaluating the platform should start with a representative customer workload. Partners should check model and framework compatibility, required versions of Huawei’s CANN software toolkit, and networking, then validate monitoring and performance under production-like loads.
VARs and infrastructure advisers should include software maturity in hardware comparisons. Migration effort and long-term support can change deployment economics even when the underlying accelerator meets a customer’s compute needs. Partners building AI infrastructure services should account for those costs when customers ask for alternatives to Nvidia.
MSPs supporting mixed accelerator environments should keep platform-specific dependencies visible. Providers of managed AI infrastructure may need separate runbooks and escalation paths when CUDA and CANN coexist in the same customer estate.
For China-based partners, broader software support could make Ascend easier to package with integration and managed services rather than as a hardware-only sale. More of the value around those projects could come from the services attached to deployment and ongoing support.
More APAC news: One-person startups are gaining ground in China as AI takes on coding, marketing, customer support, and other early-stage work.




