DeepSeek and Huawei Release Open-Source AI Programming Suite to Challenge Nvidia's CUDA
Developers now have a direct, open-source pathway to program AI models on Huawei hardware without relying on Nvidia's proprietary software ecosystem. The joint release of Ascend-optimized libraries by DeepSeek and Huawei aims to break the industry's deep dependence on CUDA.
By Naina Verma
AI developers building large language models can now compile and run their code directly on Huawei's Ascend processors using a newly released open-source software stack. The joint release from DeepSeek and Huawei bypasses Nvidia's proprietary CUDA platform entirely, offering a functional alternative for AI training.[1][2]
The companies have published Ascend-optimized versions of three critical infrastructure libraries: TileLang, DeepGEMM, and DeepEP. These tools provide the underlying compute and communication frameworks necessary to train and deploy massive AI models on non-Nvidia silicon.[4][6]
Nvidia's dominance in the artificial intelligence sector stems largely from CUDA, a closed-source programming model that developers have spent over a decade learning. Breaking that monopoly requires not just competitive hardware, but a software ecosystem that developers actually want to use.[3][4]
While marketing materials often frame new frameworks as universal CUDA killers, this specific release is highly targeted. It is engineered explicitly to make Huawei's Ascend chips viable for the complex workloads demanded by frontier AI models, rather than serving as a drop-in replacement for all hardware.[1][5]
The Technical Stack
TileLang serves as the foundation of the new release, simplifying the creation of custom computational kernels. Writing these kernels traditionally required deep, hardware-specific knowledge that kept developers tethered to Nvidia's well-documented ecosystem.[2][4]
By open-sourcing Ascend support for TileLang, DeepSeek has lowered the barrier to entry for programming Huawei's silicon. Developers can now write high-performance code without needing to master the proprietary intricacies of the Ascend architecture from scratch.[4][6]
DeepGEMM addresses the core mathematical workload of artificial intelligence: matrix multiplication. The library is optimized to handle the dense calculations required by large language models, ensuring that Huawei's chips are utilized efficiently during training runs.[6]
DeepEP tackles the logistical challenge of distributed computing. When a model is too large to fit on a single processor, DeepEP manages the complex, high-speed communication required to split the workload across thousands of interconnected chips.[2][6]
"The tools include compute and communication libraries, as well as Ascend support for TileLang," Tom's Hardware reported, noting that the suite is specifically designed to reduce reliance on the established Nvidia ecosystem.[2]
Hardware and Geopolitics
The software rollout coincides with Huawei detailing its SuperPoD Flex architecture. This scalable cluster design allows data center operators to link thousands of Ascend processors together to train massive models efficiently.[6]
Pairing open-source software with scalable hardware represents a direct attempt to replicate the vertical integration that made Nvidia a multi-trillion-dollar company. Huawei is providing the silicon, while DeepSeek is providing the translation layer for developers.[1][3]
This development carries significant weight for regions facing export restrictions on advanced American semiconductors. With access to Nvidia's latest H100 and Blackwell chips heavily restricted in China, domestic tech giants have been forced to build their own infrastructure.[1][3]
DeepSeek, an AI research company that has gained international traction for its highly efficient models, brings crucial credibility to the project. Their endorsement and active development of the Ascend software stack signals to other developers that the hardware is production-ready.[4][5]
The Adoption Challenge
Releasing open-source code is only the first step in challenging a deeply entrenched software monopoly. Nvidia's CUDA benefits from a massive, compounding network effect: because everyone uses it, all new AI research defaults to it.[3][5]
To succeed, the DeepSeek and Huawei alliance must convince independent developers and rival tech companies to invest time in learning the Ascend ecosystem. A software stack is only as valuable as the community actively maintaining and building upon it.[1][4]
Early indicators suggest the tools are functional, but transitioning existing codebases away from CUDA remains a labor-intensive process. Companies must weigh the cost of rewriting their infrastructure against the strategic benefit of hardware independence.[3][5]
The open-source nature of the release is a calculated strategy to accelerate this adoption. By allowing developers to inspect, modify, and distribute the code freely, Huawei and DeepSeek are attempting to crowdsource the optimization of their platform.[4][6]
Looking Ahead
The true test of this software suite will come over the next twelve months, as third-party data centers attempt to deploy it at scale. Success will be measured not by GitHub stars, but by the number of frontier models trained exclusively on Ascend hardware.[1][5]
Key points
- DeepSeek and Huawei have released open-source Ascend-optimized libraries, including TileLang, DeepGEMM, and DeepEP.
- The software suite allows developers to compile and run large language models on Huawei silicon without Nvidia's CUDA.
- The release coincides with Huawei detailing its SuperPoD Flex architecture for scalable AI data centers.
- While functional, the tools face the massive challenge of overcoming CUDA's entrenched decade-long network effect.
What we don’t know
- It remains unclear how much performance degradation occurs when porting existing CUDA-optimized models to the Ascend stack.
- The exact adoption rate among independent, non-Chinese AI developers has not yet been measured.
- Neither company has disclosed the financial investment required to maintain and update these open-source libraries long-term.
How we got here
2006
Nvidia releases the first version of CUDA, establishing the foundation for GPU-accelerated computing.
2023
US export controls heavily restrict the sale of advanced Nvidia AI chips to Chinese technology companies.
Oct 2026
DeepSeek and Huawei jointly release open-source Ascend programming tools to bypass the CUDA ecosystem.
- Open-Source Advocates
- Argue that open frameworks are essential to break hardware monopolies and lower AI development costs.
- Incumbent Ecosystem Defenders
- Maintain that CUDA's decade-long head start and deep integration make switching costs prohibitively high.
- Strategic Autonomy Analysts
- Focus on the geopolitical necessity of developing domestic AI stacks to bypass export controls.
Perspectives this story doesn't cover
- Independent AI developers tasked with migrating codebases
- Nvidia software engineers
Sources
[1]South China Morning PostStrategic Autonomy AnalystsChina's DeepSeek open-sources tools to help Huawei chips supplant Nvidia in AI
Read on South China Morning Post →
[2]Tom's HardwareIncumbent Ecosystem DefendersDeepSeek and Huawei release open-source Ascend AI programming tools to reduce reliance on Nvidia CUDA ecosystem
Read on Tom's Hardware →
[3]QuartzStrategic Autonomy AnalystsDeepSeek and Huawei are partnering to build open-source AI chip software to cut Nvidia reliance
Read on Quartz →
[4]The Next WebOpen-Source AdvocatesDeepSeek open-sources Huawei chip tools as a simpler alternative to CUDA
Read on The Next Web →
[5]TechzineIncumbent Ecosystem DefendersDeepSeek brings AI software to Huawei's Ascend chips
Read on Techzine →
[6]PandailyOpen-Source AdvocatesDeepSeek Open-Sources Ascend Versions of TileLang, DeepGEMM and DeepEP as Huawei Details SuperPoD Flex
Read on Pandaily →
More in Technology
See all →Open-Source Funding
The Four Models of Open-Source Sustainability: How Free Software Funds Its Own Survival
8 sources
Institutional Blockchain
Linux Foundation Decentralized Trust Adds Swift and Wells Fargo, Signaling Open-Source Finance Infrastructure Shift
7 sources
GitLab Security
GitLab Email Token Flaw Allows Code Push to Protected Repositories
5 sources
Agent Security
Critical 'Plugin4Shell' Flaw Allows Zero-Click RCE in GitHub Copilot and Other AI Coding Agents
6 sources
Comments
Every angle. Every day.
Get Technology stories with full source coverage and perspective breakdowns, free every day.




