rkj dev

DeepSeek and Huawei Release Open-Source Ascend AI Tools

The release brings TileLang compiler support and the TileKernels library to Huawei Ascend 950 chips, targeting cross-hardware flexibility.

An engineer working in a modern server facility with compute clusters
Illustration: DeepSeek and Huawei released open-source programming tools for Huawei's Ascend AI chips.AI-generated illustration

Key takeaways

  • DeepSeek and Huawei open-sourced programming tools and libraries for Ascend 950 AI chips on September 30, 2026.
  • TileLang officially added native code generation, automatic scheduling, and synchronization for Huawei Ascend 950 NPUs.
  • DeepSeek updated its open-source TileKernels library with an Ascend backend, allowing the same Python APIs to run across Nvidia GPUs and Huawei NPUs.
  • The engineering effort included computation and communication optimization across a supernode system of 128 Ascend 950 accelerators.

On September 30, 2026, Chinese artificial intelligence laboratory DeepSeek and hardware vendor Huawei released open-source programming tools and libraries designed for Huawei's Ascend AI accelerators. As reported by Tom's Hardware via Reuters, the toolchain aims to ease development on Huawei hardware and reduce software dependence on Nvidia's CUDA ecosystem.

The update integrates Ascend support into TileLang, a domain-specific programming language and compiler infrastructure, alongside a multi-backend update to DeepSeek's TileKernels operator library.

An AI accelerator processor integrated onto a circuit board
Illustration: TileLang's Ascend 950 backend provides native code generation for Huawei NPUs.AI-generated illustration

What the New Ascend Software Toolchain Includes

The software release introduces native compilation and pre-built kernels for neural processing units (NPUs). According to the TileLang repository, the project officially added an Ascend 950 backend on September 30, 2026. This backend provides native code generation, automatic scheduling and synchronization, and SIMD/SIMT vector programming for the accelerator, built atop Apache TVM compiler infrastructure.

Simultaneously, DeepSeek updated its TileKernels repository, an MIT-licensed library of optimized operations used in the company's internal model training and inference workloads. With the September 30 update, TileKernels introduced runtime hardware detection that allows identical Python APIs to run across both Nvidia GPUs and Huawei NPUs.

According to the repository documentation, the library covers key large language model operations, including:

  • Mixture-of-Experts (MoE) routing with top-k expert selection
  • Quantization routines, including per-token, per-block, and per-channel FP8 and FP4 casting and dequantization
  • Engram gating kernels with fused RMSNorm and autograd wrappers
  • Manifold HyperConnection (mHC) kernels, including Sinkhorn normalization
  • Rotary position embedding (RoPE) transform kernels
Software researchers collaborating in a development laboratory
Illustration: Joint hardware and software optimization helps drive multi-accelerator compute clusters.AI-generated illustration

Hardware Requirements and Cluster Optimization

To run the Ascend backend, the TileKernels project specification mandates Python 3.12 or higher, PyTorch 2.13 or higher, TileLang 0.1.15 or higher, Huawei Ascend 950 NPUs, and Huawei CANN 9.2.0 or higher. For its parallel CUDA path, the library requires NVIDIA SM90 or SM100 architecture GPUs and CUDA Toolkit 13.1 or higher.

Beyond individual chip operations, the collaboration targeted multi-accelerator data transmission. As reported by Tom's Hardware, DeepSeek and Huawei worked together to optimize computation and chip-to-chip communication on a supernode system built around 128 Ascend 950 chips. This system design focuses on accelerating local calculations while maintaining high data-transfer throughput across the interconnect to keep processors utilized during distributed training and inference.

Challenging the CUDA Ecosystem

Software programmability has historically been a significant barrier to alternative AI hardware adoption. As highlighted by The Daily Upside, competing toolchains such as AMD's ROCm, Intel's oneAPI, and the multi-vendor OpenCL framework have spent years attempting to replicate CUDA's deep developer adoption and extensive kernel support.

According to The Daily Upside, DeepSeek stated that TileLang is designed to be simpler to program than CUDA while continuing to drive hardware near its upper operational limits. Huawei also indicated plans to release next-generation Ascend chips in early 2027, with TileLang positioned to program the upcoming hardware.

Frequently asked questions

What chips are supported by the new TileKernels Ascend backend?

The Ascend backend in TileKernels requires Huawei Ascend 950 NPUs and the CANN 9.2.0 or higher software stack.

Is TileKernels open source?

Yes, DeepSeek has published TileKernels on GitHub under the open-source MIT License.

What operators are included in the open-source release?

The library includes kernels for Mixture of Experts (MoE) routing, FP8/FP4 quantization, Engram gating, Manifold HyperConnections, and Rotary Position Embeddings (RoPE).

Sources

  1. deepseek-ai/TileKernelsGitHub · Apr 22, 2026 · Official
  2. tile-ai/tilelangGitHub · Oct 3, 2024 · Official
  3. DeepSeek and Huawei release open-source Ascend AI programming tools to reduce reliance on Nvidia CUDA ecosystem — tools include compute and communication libraries, as well as Ascend support fo…Tom's Hardware · Oct 1, 2026
  4. DeepSeek, Huawei Challenge the Software Powering Nvidia’s AI DominanceThe Daily Upside · Oct 1, 2026

How this story was made: the newsroom picked it up from Google Search, Google News and tomshardware.com, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (33 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.

#DeepSeek #Huawei #Ascend 950 #TileLang #AI Chips #Open Source

Published October 2, 2026 at 00:43 UTC