Qualcomm Unveils 'Dragonfly' Chip Architecture, Challenging Nvidia's Dominance by Eliminating High-Bandwidth Memory
Qualcomm has introduced a radically new data center architecture that physically stacks low-power memory directly onto processors, bypassing the industry's reliance on scarce, power-hungry High-Bandwidth Memory (HBM). Backed by anchor agreements with Meta and Microsoft, the 'Dragonfly' platform aims to slash the energy and financial costs of AI inference.
By Logan Price
- AI Infrastructure Analysts
- Focusing on total cost of ownership and the transition from AI training to AI inference.
- Semiconductor Supply Chain
- Emphasizing the strategic bypass of global packaging bottlenecks.
- Hyperscale Cloud Operators
- Valuing multi-vendor flexibility and the disruption of software monopolies.
Perspectives this story doesn't cover
- Environmental Advocates
- Incumbent GPU Manufacturers
The generative AI boom has hit a physical wall: memory. As large language models grow more complex, the industry has become entirely dependent on High-Bandwidth Memory (HBM)—a specialized, expensive, and power-hungry component that sits next to the processor. Moving data laterally between the compute die and the HBM stack across a silicon interposer now consumes a massive portion of a data center's energy budget.
On Wednesday, Qualcomm unveiled a radical alternative designed to break this bottleneck. During its Investor Day in New York, the company introduced "Dragonfly," a comprehensive new data center portfolio that fundamentally reimagines how AI chips are built. Instead of relying on traditional HBM, Qualcomm is introducing a proprietary architecture called High-Bandwidth Compute (HBC).[1]
The HBC architecture physically stacks low-power DRAM (LPDDR) directly on top of the logic processor using 3D silicon bonding, specifically utilizing Through-Silicon Vias (TSVs). By placing the memory vertically above the compute cores rather than beside them, Qualcomm effectively collapses the distance data must travel to near zero.[1][2]
This vertical integration eliminates the need for the expensive 2.5D silicon interposers that currently bottleneck the production of Nvidia's flagship GPUs. More importantly, it leverages the highly efficient LPDDR memory technology that Qualcomm has spent the last fifteen years perfecting for the mobile smartphone market, scaling it up for enterprise data centers.[3]
The efficiency gains are striking. Qualcomm claims the HBC architecture delivers up to six times higher memory bandwidth per watt compared to current HBM-based designs. By mitigating the thermal output and power consumption associated with lateral data movement, the company projects its new accelerators will offer four to eight times better performance-per-watt than incumbent GPU systems.[2]
The flagship of this new portfolio is the Dragonfly C1000 CPU. Built on Qualcomm's custom Oryon architecture—developed following its acquisition of Nuvia—the processor integrates more than 250 cores on a single chiplet design. Operating at sustained frequencies above 5GHz, the C1000 is engineered specifically for the dense, continuous reasoning demands of "agentic" AI workloads.[1]
The flagship of this new portfolio is the Dragonfly C1000 CPU.
Alongside the CPU, Qualcomm detailed its upcoming AI inference accelerators. The AI250, slated for commercial sampling in mid-2027, will feature the first generation of HBC technology, delivering an effective memory bandwidth of 133 terabytes per second. Its successor, the AI300, promises to push those boundaries even further to handle massive multimodal models.[2]
Hardware alone cannot unseat Nvidia, whose CUDA software platform remains the industry standard. To bridge this gap, Qualcomm confirmed its $3.92 billion all-stock acquisition of Modular, an AI software startup founded by Chris Lattner, the creator of Apple's Swift programming language.
Modular's MAX inference engine and Mojo programming language allow developers to write AI code once and deploy it seamlessly across CPUs, GPUs, and custom silicon. By commoditizing the hardware layer, Qualcomm hopes to lower the switching costs that currently keep developers locked into the Nvidia ecosystem.[2]
Hyperscale cloud providers, eager to reduce their reliance on a single dominant supplier, are already signing on. Meta CEO Mark Zuckerberg confirmed a multi-generation agreement to deploy the Dragonfly C1000 in Meta's next-generation server fleet. Microsoft has similarly committed to deploying Qualcomm's HBC architecture within its Azure cloud infrastructure.
These anchor partnerships lend immediate credibility to Qualcomm's 2028 delivery timeline. They also highlight a broader industry shift: as AI matures from the training phase—where Nvidia's raw computational power is unmatched—to the inference phase, where models are deployed to millions of users, efficiency and total cost of ownership become the defining metrics.[2][3]
The supply chain implications are equally significant. Because HBC utilizes standard LPDDR memory and avoids complex interposer packaging, Qualcomm sidesteps the severe capacity constraints at TSMC's CoWoS (Chip-on-Wafer-on-Substrate) packaging facilities, which have historically limited the supply of high-end AI accelerators.[1]
Qualcomm is projecting $15 billion in data center revenue by fiscal 2029, a massive expansion from its traditional base in mobile telecommunications. If the Dragonfly architecture scales as promised, it could fundamentally democratize AI infrastructure, allowing data centers to run vastly more capable models without requiring a proportional increase in power consumption.
Key points
- Qualcomm unveiled its 'Dragonfly' data center portfolio, featuring the new High-Bandwidth Compute (HBC) architecture.
- HBC stacks low-power memory directly onto the processor, eliminating the need for expensive High-Bandwidth Memory (HBM).
- The architecture promises up to six times higher memory bandwidth per watt compared to current GPU systems.
- Meta and Microsoft have signed anchor agreements to deploy Qualcomm's new chips in their data centers.
- Qualcomm acquired AI software startup Modular for $3.92 billion to challenge Nvidia's CUDA software monopoly.
Why this matters
By eliminating the most expensive and power-hungry bottleneck in AI hardware, Qualcomm's new architecture promises to drastically lower the cost of running advanced AI models. This shift could democratize access to enterprise-grade AI, making smart agents cheaper and more ubiquitous across everyday applications while slashing the massive energy footprint of modern data centers.
What we don’t know
- Whether the 3D-stacked memory architecture can effectively manage the intense thermal output of 250-core processors at maximum rack density.
- How quickly developers will adopt Modular's software stack over the deeply entrenched Nvidia CUDA ecosystem.
- Whether incumbent GPU manufacturers will accelerate their own near-memory compute solutions before Dragonfly's 2028 rollout.
Key terms
- High-Bandwidth Memory (HBM)
- A specialized, high-performance memory technology that sits next to a processor, currently standard in top-tier AI chips but limited by high costs and power demands.
- High-Bandwidth Compute (HBC)
- Qualcomm's new architecture that stacks low-power memory directly on top of the processor to reduce data travel distance and save energy.
- Through-Silicon Via (TSV)
- A vertical electrical connection that passes completely through a silicon wafer, used to stack memory chips directly onto logic processors.
- Silicon Interposer
- An electrical interface routing between one socket or connection to another, traditionally used to connect a GPU to its adjacent HBM.
- Agentic AI
- Advanced artificial intelligence systems designed to autonomously plan, reason, and execute complex, multi-step tasks over time.
Sources
[1]The ElecSemiconductor Supply ChainQualcomm details Dragonfly C1000 and HBC technology, stacking LPDDR to replace HBM
Read on The Elec →
[2]Hyperframe ResearchAI Infrastructure AnalystsAnalyzing Qualcomm's Dragonfly: A Paradigm Shift in Agentic AI Infrastructure
Read on Hyperframe Research →
[3]VoxForAI Infrastructure AnalystsHow Qualcomm's LPDDR5 strategy lowers AI inference costs
Read on VoxFor →
Comments
More in Artificial Intelligence
See all →AI Infrastructure
How FlashAttention Bypasses the GPU Memory Bottleneck to Enable Long-Context AI
5 sources
Open Source Standards
How the Open Source Initiative's 1.0 Definition Excludes the Most Downloaded Open-Weight AI Models
7 sources
Generative Adversarial Networks
How a Generator and a Discriminator Compete to Create Realistic AI Output
8 sources
Machine Learning
How Generative AI Maps the Joint Probability Distribution of Data
5 sources
Every angle. Every day.
Get Artificial Intelligence stories with full source coverage and perspective breakdowns delivered to your inbox.




