Chinese AI company DeepSeek announced on September 30, 2026, that together with Huawei Technologies it has developed programming tools optimized for Huawei's Ascend chips. As part of the project, the source code for the high-level programming language TileLang and the related compute and communication libraries has been open-sourced. According to Reuters, the announcement was made via the company's official WeChat page. The stated goal of the initiative is to build infrastructure that serves as an alternative to Nvidia's ecosystem in AI computing. DeepSeek, itself a developer of large language models, is the first user of these tools: the company is opening to the public ecosystem the technologies it uses to train and run its models.
TileLang: A High-Level Language for Ascend
In its official post, DeepSeek presented TileLang as a high-level programming language designed for the Ascend platform. According to the company, the language offers a simpler programming model than Nvidia's CUDA platform: this increases development efficiency and simplifies code logic for AI accelerators.
"TileLang offers a simpler programming model compared to Nvidia CUDA." — DeepSeek, via its official WeChat post
Along with TileLang, the compute and communication libraries were also open-sourced. Together they form a complete programming infrastructure for Ascend chips: developers now work not with low-level hardware details but with high-level language tools. The openness of the programming infrastructure has practical significance: closed tools tie a developer to a single vendor, while open tools provide the freedom to port code and evolve it independently. This logic is exactly what underlies the move to build an open alternative to the ecosystem built around Nvidia CUDA. Partnerships in the chip industry have been multiplying lately: we previously wrote about OpenAI and Synopsys jointly developing an AI model for chip design.
DeepGEMM-Ascend: An Open-Source Kernel Library
The second result of the partnership is DeepGEMM-Ascend: an open-source library of matrix multiplication kernels for Huawei Ascend neural processing units (NPUs). Matrix multiplication is the core computation of neural networks: the speed of training and running large language models directly depends on the efficiency of these operations.
The first release came out on September 30, 2026, and supports Ascend 950 devices. According to DeepSeek, the kernels run at near-hardware-limit performance. The library is fully compatible with the DeepGEMM API: GEMM operations in BF16, FP8, and FP4 formats, as well as MQA logits and MegaMoE kernels, run on the Ascend platform. The project is openly hosted on GitHub — anyone can review the code, test it, and use it in their own projects. Achieving near-hardware-limit performance directly affects the cost of running large models: more computation per chip means less energy and less hardware. So the efficiency of kernel libraries is not only a matter of speed, but also of economics.
A "Supernode" Based on 128 Ascend 950 Chips
The two companies have also jointly promoted a "supernode" solution based on 128 Ascend 950 chips. It optimizes compute power and inter-chip communication simultaneously — these two factors are the main bottleneck when training large models. According to DeepSeek, Huawei fully supported the infrastructure development. This shows the depth of the partnership: it is not just about a single library, but about co-designing hardware and software.
The supernode architecture combines individual chips into a single powerful compute node. This approach is widely used in data centers: instead of separate processors, their dense cluster operates as one logical system. The DeepSeek and Huawei solution aims to standardize such a cluster specifically for the Ascend ecosystem — and because the programming tools are open, third-party developers can also build solutions for this architecture.




