DeepSeek and Huawei open-source TileLang and Ascend 950 libraries as a CUDA alternative

NVIDIA's strongest moat has never been only its chips. It is CUDA, the programming ecosystem that nearly all AI software is written for. DeepSeek and Huawei are now attacking that moat directly. They open-sourced Ascend support for TileLang, a framework for writing high-performance AI kernels, along with six modules including DeepGEMM-Ascend and DeepEP-Ascend. These ports bring DeepSeek's well-regarded matrix multiplication and expert-parallel communication libraries, originally tuned for NVIDIA hardware, to Huawei's Ascend 950 processors.
The technical aim is to make Ascend clusters a realistic target for frontier training and inference. The libraries support 128-card supernodes, Huawei's large interconnected cluster configuration, and Huawei says Ascend 950 supernodes fully support DeepSeek V4. DeepSeek has already used Ascend hardware to train its V4-Flash models and plans an Inner Mongolia facility with 160,000 Ascend accelerators. Shipping these tools as open source means other Chinese labs can use them, which strengthens the whole domestic ecosystem instead of just DeepSeek.
This is part of a broader push toward sovereign compute in China. Alibaba unveiled its own Zhenwu V900 processor at Apsara in the same week, and practitioners on r/LocalLLaMA are already posting benchmarks of Qwen3.8-Flash-Next running on two 96GB Ascend cards with vLLM. Export controls have pushed Chinese labs toward this path. The open question has been whether software would hold back domestic chips more than the hardware itself. A credible TileLang port narrows that gap.
The caveats: CUDA has more than 15 years of tooling, libraries and developer habits behind it, and one framework port does not replace that. Performance numbers against comparable NVIDIA setups are still scarce. DeepSeek users also reported its servers being busy or down again this week, a reminder that capacity is still tight. Watch for third-party benchmarks of Ascend 950 versus Blackwell on DeepSeek V4 and for adoption by labs beyond DeepSeek.