Skip to main content
AI InfrastructureHuawei· 4 min read· in Technology

DeepSeek and Huawei Release Open-Source AI Programming Suite to Challenge Nvidia's CUDA

Developers now have a direct, open-source pathway to program AI models on Huawei hardware without relying on Nvidia's proprietary software ecosystem. The joint release of Ascend-optimized libraries by DeepSeek and Huawei aims to break the industry's deep dependence on CUDA.

By Naina Verma

AI developers building large language models can now compile and run their code directly on Huawei's Ascend processors using a newly released open-source software stack. The joint release from DeepSeek and Huawei bypasses Nvidia's proprietary CUDA platform entirely, offering a functional alternative for AI training.[1][2]

The companies have published Ascend-optimized versions of three critical infrastructure libraries: TileLang, DeepGEMM, and DeepEP. These tools provide the underlying compute and communication frameworks necessary to train and deploy massive AI models on non-Nvidia silicon.[4][6]

Nvidia's dominance in the artificial intelligence sector stems largely from CUDA, a closed-source programming model that developers have spent over a decade learning. Breaking that monopoly requires not just competitive hardware, but a software ecosystem that developers actually want to use.[3][4]

While marketing materials often frame new frameworks as universal CUDA killers, this specific release is highly targeted. It is engineered explicitly to make Huawei's Ascend chips viable for the complex workloads demanded by frontier AI models, rather than serving as a drop-in replacement for all hardware.[1][5]

The new open-source libraries map directly to functions traditionally handled by Nvidia's CUDA platform.

The Technical Stack

TileLang serves as the foundation of the new release, simplifying the creation of custom computational kernels. Writing these kernels traditionally required deep, hardware-specific knowledge that kept developers tethered to Nvidia's well-documented ecosystem.[2][4]

By open-sourcing Ascend support for TileLang, DeepSeek has lowered the barrier to entry for programming Huawei's silicon. Developers can now write high-performance code without needing to master the proprietary intricacies of the Ascend architecture from scratch.[4][6]

DeepGEMM addresses the core mathematical workload of artificial intelligence: matrix multiplication. The library is optimized to handle the dense calculations required by large language models, ensuring that Huawei's chips are utilized efficiently during training runs.[6]

DeepEP tackles the logistical challenge of distributed computing. When a model is too large to fit on a single processor, DeepEP manages the complex, high-speed communication required to split the workload across thousands of interconnected chips.[2][6]

"The tools include compute and communication libraries, as well as Ascend support for TileLang," Tom's Hardware reported, noting that the suite is specifically designed to reduce reliance on the established Nvidia ecosystem.[2]

Hardware and Geopolitics

The software rollout coincides with Huawei detailing its SuperPoD Flex architecture. This scalable cluster design allows data center operators to link thousands of Ascend processors together to train massive models efficiently.[6]

Illustration: Transitioning existing AI models to a new software stack requires developers to rewrite foundational compute kernels.

Pairing open-source software with scalable hardware represents a direct attempt to replicate the vertical integration that made Nvidia a multi-trillion-dollar company. Huawei is providing the silicon, while DeepSeek is providing the translation layer for developers.[1][3]

This development carries significant weight for regions facing export restrictions on advanced American semiconductors. With access to Nvidia's latest H100 and Blackwell chips heavily restricted in China, domestic tech giants have been forced to build their own infrastructure.[1][3]

DeepSeek, an AI research company that has gained international traction for its highly efficient models, brings crucial credibility to the project. Their endorsement and active development of the Ascend software stack signals to other developers that the hardware is production-ready.[4][5]

The Adoption Challenge

Releasing open-source code is only the first step in challenging a deeply entrenched software monopoly. Nvidia's CUDA benefits from a massive, compounding network effect: because everyone uses it, all new AI research defaults to it.[3][5]

To succeed, the DeepSeek and Huawei alliance must convince independent developers and rival tech companies to invest time in learning the Ascend ecosystem. A software stack is only as valuable as the community actively maintaining and building upon it.[1][4]

Early indicators suggest the tools are functional, but transitioning existing codebases away from CUDA remains a labor-intensive process. Companies must weigh the cost of rewriting their infrastructure against the strategic benefit of hardware independence.[3][5]

Alternative hardware ecosystems are expanding as companies seek to reduce their reliance on a single silicon vendor.

The open-source nature of the release is a calculated strategy to accelerate this adoption. By allowing developers to inspect, modify, and distribute the code freely, Huawei and DeepSeek are attempting to crowdsource the optimization of their platform.[4][6]

Looking Ahead

The true test of this software suite will come over the next twelve months, as third-party data centers attempt to deploy it at scale. Success will be measured not by GitHub stars, but by the number of frontier models trained exclusively on Ascend hardware.[1][5]

If the ecosystem matures, it could fracture the global AI infrastructure market into two distinct camps: one reliant on Nvidia and CUDA, and another built on open-source tools and alternative silicon.[3][4]

For now, developers have a tangible, shipped alternative to evaluate. The monopoly on AI programming software has not been broken, but the foundation for a viable parallel ecosystem has officially been laid.[1][2]

Key points

  1. DeepSeek and Huawei have released open-source Ascend-optimized libraries, including TileLang, DeepGEMM, and DeepEP.
  2. The software suite allows developers to compile and run large language models on Huawei silicon without Nvidia's CUDA.
  3. The release coincides with Huawei detailing its SuperPoD Flex architecture for scalable AI data centers.
  4. While functional, the tools face the massive challenge of overcoming CUDA's entrenched decade-long network effect.

What we don’t know

  • It remains unclear how much performance degradation occurs when porting existing CUDA-optimized models to the Ascend stack.
  • The exact adoption rate among independent, non-Chinese AI developers has not yet been measured.
  • Neither company has disclosed the financial investment required to maintain and update these open-source libraries long-term.

How we got here

  1. 2006

    Nvidia releases the first version of CUDA, establishing the foundation for GPU-accelerated computing.

  2. 2023

    US export controls heavily restrict the sale of advanced Nvidia AI chips to Chinese technology companies.

  3. Oct 2026

    DeepSeek and Huawei jointly release open-source Ascend programming tools to bypass the CUDA ecosystem.

Open-Source Advocates 35%Incumbent Ecosystem Defenders 35%Strategic Autonomy Analysts 30%
Open-Source Advocates
Argue that open frameworks are essential to break hardware monopolies and lower AI development costs.
Incumbent Ecosystem Defenders
Maintain that CUDA's decade-long head start and deep integration make switching costs prohibitively high.
Strategic Autonomy Analysts
Focus on the geopolitical necessity of developing domestic AI stacks to bypass export controls.

Perspectives this story doesn't cover

  • Independent AI developers tasked with migrating codebases
  • Nvidia software engineers

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Open-Source Advocates 35%Incumbent Ecosystem Defenders 35%Strategic Autonomy Analysts 30%
  1. [1]South China Morning PostStrategic Autonomy Analysts

    China's DeepSeek open-sources tools to help Huawei chips supplant Nvidia in AI

    Read on South China Morning Post →
  2. [2]Tom's HardwareIncumbent Ecosystem Defenders

    DeepSeek and Huawei release open-source Ascend AI programming tools to reduce reliance on Nvidia CUDA ecosystem

    Read on Tom's Hardware →
  3. [3]QuartzStrategic Autonomy Analysts

    DeepSeek and Huawei are partnering to build open-source AI chip software to cut Nvidia reliance

    Read on Quartz →
  4. [4]The Next WebOpen-Source Advocates

    DeepSeek open-sources Huawei chip tools as a simpler alternative to CUDA

    Read on The Next Web →
  5. [5]TechzineIncumbent Ecosystem Defenders

    DeepSeek brings AI software to Huawei's Ascend chips

    Read on Techzine →
  6. [6]PandailyOpen-Source Advocates

    DeepSeek Open-Sources Ascend Versions of TileLang, DeepGEMM and DeepEP as Huawei Details SuperPoD Flex

    Read on Pandaily →

Comments

Stay informed

Every angle. Every day.

Get Technology stories with full source coverage and perspective breakdowns, free every day.