OpenAI Unveils 'Jalapeño' Inference Chip With Broadcom to Control Full Stack and Cut Compute Costs
OpenAI has partnered with Broadcom to launch its first custom silicon accelerator, designed specifically to slash the skyrocketing costs of running large language models in production.
By Factlen Editorial Team
- OpenAI & Hardware Partners
- Custom silicon is the only sustainable path to scaling frontier AI.
- Financial Analysts
- Inference efficiency is the key to OpenAI's valuation and eventual IPO.
- Merchant Silicon Incumbents
- Custom chips serve a niche, but GPUs remain the undisputed engine of AI progress.
What's not represented
- · Smaller AI startups unable to afford custom silicon
- · Energy grid operators managing data center power loads
Why this matters
As AI usage scales globally, the ongoing cost of generating answers—known as inference—has become the industry's biggest financial bottleneck. By designing its own silicon, OpenAI aims to drastically lower the price of AI access while securing its path to profitability ahead of a highly anticipated public offering.
Key points
- OpenAI and Broadcom have unveiled 'Jalapeño,' a custom silicon chip designed exclusively for AI inference.
- The chip strips away training architecture to focus on data movement, aiming to cut inference costs by roughly 50%.
- Inference currently accounts for 80% to 90% of the lifetime computing costs for production AI systems.
- OpenAI plans to deploy the chips by the end of 2026 to improve margins ahead of a rumored public offering.
OpenAI has officially entered the semiconductor arena, partnering with Broadcom to unveil "Jalapeño," the artificial intelligence company’s first custom-designed silicon accelerator. Billed as an "Intelligence Processor," the new chip represents a major strategic shift for the ChatGPT creator, moving the company beyond software and model development into the physical infrastructure that powers its products. The announcement, which also includes manufacturing partner Celestica, outlines a multi-generation compute platform designed to make advanced AI faster, more reliable, and significantly cheaper to operate at a global scale.[1][2]
Unlike the versatile graphics processing units (GPUs) that currently dominate the AI hardware market, Jalapeño is an application-specific integrated circuit (ASIC) built from a blank slate. It is engineered for one specific task: large language model inference. While general-purpose chips must balance the complex mathematics required to train new models with the speed needed to serve them, Jalapeño strips away the training architecture entirely. Instead, it focuses on the precise bottlenecks that slow down live AI products, optimizing for rapid data movement, memory bandwidth, and low-latency networking.[1]
To understand the significance of the move, one must untangle the two distinct phases of artificial intelligence: training and inference. Training is the highly publicized, computationally massive process of teaching a model how to think by feeding it trillions of words over several months. Inference, by contrast, is the ongoing, everyday process of running live user prompts through that finished model to generate a response. While training requires a massive upfront investment, inference is a perpetual operational expense that scales directly with user adoption.[4]

The economic reality of the AI boom is that inference has become the industry's heaviest financial anchor. According to recent infrastructure analyses, inference workloads are projected to consume 65% of all AI compute by 2029. More critically, because inference runs continuously to serve millions of daily queries, it now accounts for 80% to 90% of the lifetime cost of a production AI system. Every time a user generates a document, writes a line of code, or asks a chatbot a question, a meter is running in a data center.
For OpenAI, which serves hundreds of millions of active users and powers thousands of enterprise applications through its API, those running meters translate to exorbitant cloud bills. The broader industry is feeling the same squeeze; a recent developer survey found that 44% of organizations now allocate over three-quarters of their total AI budgets strictly to inference. By designing hardware specifically tuned to the quirks of its own models, OpenAI hopes to break this cost curve and make high-volume AI deployment financially sustainable.[4]
The hardware architecture of Jalapeño reflects a deep collaboration between OpenAI's software engineers and Broadcom's silicon designers. Because OpenAI intimately understands the specific serving patterns, memory movement, and software kernels of its own frontier models, it was able to design a chip that eliminates the inefficiencies found in off-the-shelf hardware. Broadcom provided the foundational silicon expertise, including integrating its high-performance Tomahawk networking chips to ensure that thousands of Jalapeño processors can communicate seamlessly across massive data center racks.
The hardware architecture of Jalapeño reflects a deep collaboration between OpenAI's software engineers and Broadcom's silicon designers.
This level of vertical integration—controlling the entire technology stack from the user interface down to the physical silicon—has long been the holy grail for major technology companies. By operating across the full stack, OpenAI can co-design its future models and its future chips simultaneously. If a new AI architecture requires a specific type of memory access, the hardware team can build that capability directly into the next generation of Jalapeño, creating a feedback loop that off-the-shelf hardware buyers simply cannot match.[1][2]
While OpenAI and Broadcom have kept exact performance benchmarks under wraps pending a full technical paper, the financial targets are aggressive. Early reports indicate that Jalapeño is designed to cut OpenAI's inference costs by roughly 50%. Furthermore, the companies claim the chip will deliver significantly higher performance per watt than current leading-edge hardware, a crucial metric as data centers worldwide strain against the limits of local electrical grids.[2]

The development timeline for Jalapeño has set a blistering pace for the notoriously slow semiconductor industry. The project moved from early schematics to physical engineering samples in just nine months—a cycle that typically takes years. OpenAI attributed this speed to a software-hardware co-development process that actively utilized the company's own AI models to accelerate parts of the chip design, effectively using artificial intelligence to build the infrastructure for the next generation of artificial intelligence.[2]
The chips are already moving out of the theoretical phase. Engineering samples of Jalapeño are currently operating in OpenAI's labs at target clock speeds and power levels. The company confirmed it is actively running production-grade machine learning workloads on the new silicon, including its prior-generation GPT-5.3-Codex-Spark model, proving that the blank-slate architecture can successfully handle the complex reasoning tasks required by modern agentic AI.[1]
Despite the breakthrough, the launch of Jalapeño does not signal a divorce between OpenAI and Nvidia, the undisputed king of AI hardware. Nvidia's GPUs remain the gold standard for the computationally brutal task of model training. Earlier this year, Nvidia finalized a massive direct investment into OpenAI, securing an agreement to deploy gigawatts of computing capacity using its next-generation Vera Rubin platform. Jalapeño is designed to complement, rather than replace, this infrastructure by taking over the workload once the models are fully trained.[2]

OpenAI's foray into custom silicon mirrors a broader trend among the tech industry's heaviest hitters. Google has spent a decade refining its Tensor Processing Units (TPUs), Microsoft recently escalated its bespoke silicon efforts with the Azure Maia 200 accelerator, and Meta continues to deploy its MTIA chips. As AI models become the core engine of the modern internet, relying entirely on third-party merchant silicon for daily operations is increasingly viewed as an unacceptable strategic and financial risk.[2][3]
The financial implications of Jalapeño extend far beyond the server rack. As OpenAI lays the groundwork for a highly anticipated public offering, the company faces intense pressure from private investors and public markets to demonstrate a viable path to profitability. By drastically reducing its cost of goods sold—in this case, the compute required to serve its models—OpenAI is attempting to prove that the staggering capital expenditures of the AI era will eventually yield a sustainable, high-margin business.[2]

With manufacturing partner Celestica handling the board, rack, and system-level integration, OpenAI plans to begin deploying Jalapeño across active data centers by the end of the year. If the chip performs as promised at scale, it could fundamentally alter the economics of artificial intelligence, allowing developers to deploy more complex, reasoning-heavy models to millions of users without bankrupting the companies that build them.[3]
How we got here
Late 2023
Microsoft launches Azure Maia 100, signaling the cloud provider shift toward custom AI silicon.
October 2025
OpenAI and Broadcom publicly announce their partnership to develop custom AI hardware.
February 2026
Nvidia finalizes a massive direct investment in OpenAI to secure future training infrastructure.
June 2026
OpenAI and Broadcom officially unveil the Jalapeño inference chip, moving from schematics to lab testing in just nine months.
Late 2026
Projected timeline for the initial deployment of Jalapeño chips in active data centers.
Viewpoints in depth
OpenAI & Hardware Partners
Custom silicon is the only sustainable path to scaling frontier AI.
For OpenAI and Broadcom, the Jalapeño chip is about breaking the fundamental cost curve of artificial intelligence. They argue that relying on general-purpose GPUs for inference is inherently inefficient, as those chips carry expensive architecture designed for training that goes unused during live deployment. By controlling the full stack—from the model's software kernels down to the physical networking—they believe they can achieve a level of hardware utilization and cost-efficiency that merchant silicon simply cannot match, ultimately democratizing access to advanced reasoning models.
Financial Analysts
Inference efficiency is the key to OpenAI's valuation and eventual IPO.
Market watchers view the Jalapeño announcement through the lens of corporate finance. With OpenAI reportedly laying the groundwork for a public offering, the company must prove it can transition from a cash-burning research lab into a high-margin software business. Analysts note that because inference accounts for up to 90% of a model's lifetime cost, cutting those expenses by 50% directly transforms the company's unit economics. For investors, custom silicon is less about technical elegance and more about proving that the AI business model is fundamentally profitable.
Merchant Silicon Incumbents
Custom chips serve a niche, but GPUs remain the undisputed engine of AI progress.
While acknowledging the role of custom inference chips, defenders of merchant silicon—most notably Nvidia—maintain that general-purpose GPUs will continue to dominate the broader AI landscape. They point out that AI models are evolving at a breakneck pace, and hard-coding an ASIC for today's architectures risks obsolescence if the underlying math changes tomorrow. Furthermore, they emphasize that the most computationally intensive phase of AI—training the next generation of frontier models—still strictly requires the massive parallel processing power that only advanced GPUs can provide.
What we don't know
- The exact performance benchmarks and power efficiency metrics of the Jalapeño chip compared to Nvidia's latest inference hardware.
- Whether OpenAI intends to lease Jalapeño compute capacity to external enterprise customers, or reserve it exclusively for its own models.
- How quickly Celestica and Broadcom can scale manufacturing to meet OpenAI's massive global data center footprint.
Key terms
- Inference
- The phase of artificial intelligence where a trained model processes new data or user prompts to generate a prediction, decision, or text response.
- ASIC (Application-Specific Integrated Circuit)
- A microchip designed for a single, specific purpose—such as running AI inference—rather than general-purpose computing.
- Full-Stack Integration
- A strategy where a single company controls every layer of a technology product, from the user-facing software down to the physical hardware and networking.
- AI Training
- The initial, highly compute-intensive process of feeding massive datasets into a neural network so it can learn patterns and behaviors.
- Merchant Silicon
- Microchips designed and sold by independent semiconductor companies (like Nvidia or AMD) for use by any customer, as opposed to custom in-house chips.
Frequently asked
What is the difference between AI training and inference?
Training is the computationally massive, one-time process of teaching an AI model how to think. Inference is the ongoing, everyday process of running live user prompts through that finished model to generate a response.
Why is OpenAI building its own chip?
Inference costs run continuously and account for up to 90% of an AI system's lifetime expense. By designing a custom chip, OpenAI can optimize for its specific models and drastically reduce its cloud computing bills.
Will OpenAI stop using Nvidia chips?
No. Nvidia's GPUs remain the industry standard for the intensive process of model training. Jalapeño is designed specifically for inference, taking over the workload only after the models are fully trained.
When will the Jalapeño chip be deployed?
Engineering samples are already running in OpenAI's labs, and the company plans to begin rolling out the processors across active data centers by the end of 2026.
Sources
[1]OpenAIOpenAI & Hardware Partners
OpenAI and Broadcom introduce Jalapeño
Read on OpenAI →[2]VentureBeatFinancial Analysts
OpenAI unveils first custom AI inference chip, Jalapeño, with Broadcom — and its development was sped-up with OpenAI's own models
Read on VentureBeat →[3]MorningstarFinancial Analysts
Broadcom Unveils First Custom Chip for OpenAI
Read on Morningstar →[4]DigitalOceanMerchant Silicon Incumbents
AI inference vs training FAQ
Read on DigitalOcean →
More in ai
See all 5 stories →AI Regulation
How 42 State Attorneys General Are Using Consumer Law to Regulate OpenAI
6 sources
Silicon Sovereignty
$1 Trillion AI Chip Selloff Follows Wave of Custom Silicon Shipments, Reshaping Compute Market
7 sources
Macroeconomics
Federal Reserve Raises US Growth Forecast, Citing Surging AI Infrastructure Investment
4 sources
Every angle. Every day.
Get ai stories with full source coverage and perspective breakdowns delivered to your inbox.







