Skip to main content
Custom SiliconIndustry ShiftAug 27, 2026, 8:26 AM· 4 min read

OpenAI Claims Custom 'Jalapeño' Chip Outperforms Nvidia GB300 in Inference Efficiency

OpenAI says its first custom AI processor, developed with Broadcom, delivers up to 1.9 times more work per watt than Nvidia's current-generation systems. The company plans to deploy the chip later this year to reduce the massive power costs of running large language models.

By Ishani Patel

OpenAI Hardware Team 35%Financial Analysts 35%Semiconductor Industry Observers 30%
OpenAI Hardware Team
Focuses on the full-stack advantage of co-designing models and silicon to maximize inference efficiency and lower operational costs.
Financial Analysts
Highlights the tension between OpenAI's custom chip success and its continued financial and operational reliance on Nvidia for training hardware.
Semiconductor Industry Observers
Notes the impressive engineering feat but cautions that the chip has not yet been tested against Nvidia's latest Vera Rubin architecture.

For years, the artificial intelligence industry has operated under a simple assumption: Nvidia's hardware is the undisputed king of both training and running large language models. OpenAI, the creator of ChatGPT, is one of Nvidia's largest customers, relying on the chipmaker's infrastructure to power its rapidly expanding data centers. Yet at the Hot Chips conference at Stanford University on Tuesday, OpenAI presented a direct challenge to that dominance. The company unveiled benchmark results for "Jalapeño," its first custom-built AI processor, claiming the new chip significantly outperforms Nvidia's current-generation GB300 systems in both speed and power efficiency.[1][2]

The announcement highlights a growing tension in the AI sector: while companies like OpenAI depend heavily on Nvidia, the massive electricity costs of running AI models are forcing them to develop their own specialized hardware. However, the move is not a complete break from Nvidia. OpenAI's hardware chief, Richard Ho, clarified that the company will continue to purchase Nvidia GPUs for training its models, a computationally intensive process where Nvidia remains unchallenged. Jalapeño, instead, is designed exclusively for inference—the stage where a trained model actually processes user requests and generates responses.[2][7]

Inference is where the bulk of long-term artificial intelligence costs accumulate. Every time a user asks ChatGPT a question or requests a summary, the model must retrieve data from memory and compute an answer, consuming significant power in the process. OpenAI developed Jalapeño in partnership with Broadcom to target this specific operational bottleneck. According to the company's published results on the public InferenceX benchmark, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput compared to Nvidia's GB200 and GB300 systems.[1][5]

Benchmark results published by OpenAI claim the Jalapeño chip delivers up to 1.9 times more work per watt than Nvidia's GB300.

The efficiency gains are rooted in the chip's architecture. Jalapeño is built on TSMC's N3P process and utilizes HBM4 memory, offering 15.4 terabytes per second of bandwidth. By keeping the model's state close to the processor and minimizing data movement, the chip reduces the power draw typically associated with transformer-based inference. While Jalapeño is rated at 700 watts, OpenAI reported that its sustained power consumption remained at or below 550 watts during the tested workloads, compared to Nvidia systems rated at 1,200 to 1,400 watts.[1][2][4]

Jalapeño is built on TSMC's N3P process and utilizes HBM4 memory, offering 15.4 terabytes per second of bandwidth.

Speed is the other critical factor in evaluating inference hardware, particularly for applications that require real-time user interaction. For highly interactive workloads that demand extensive back-and-forth communication, OpenAI reported that Jalapeño achieved 2.1 to 4.1 times higher performance than the Nvidia comparison systems. End-to-end latency was reduced by 1.7 to 3.6 times across the three large language models tested in the benchmark suite: OpenAI's own GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's massive 1-trillion-parameter Kimi K2.5 model.[1][5][6]

The development cycle for the Jalapeño processor was unusually rapid for custom silicon. OpenAI stated that the engineering team moved from initial design to manufacturing tape-out in just nine months. To achieve this accelerated timeline, the company utilized its own artificial intelligence models, including Codex, to write kernel code and streamline the design and verification processes. This full-stack approach—designing the models, the serving software, and the physical silicon in tandem—allows OpenAI to optimize every layer of its infrastructure based on the exact real-world workloads its data centers handle daily.[1][5][6]

The Jalapeño chip was developed in partnership with Broadcom and moved from initial design to tape-out in just nine months.

Despite the impressive benchmark numbers, industry analysts note several important caveats. Jalapeño was tested against Nvidia's GB200 and GB300 systems, but it was not compared to Nvidia's newer Vera Rubin architecture, which also utilizes HBM4 memory and is currently beginning to ship to customers. Furthermore, OpenAI normalized the benchmark results using the published power ratings of the accelerators, rather than the actual power drawn by the Nvidia chips during the tests, which could affect the exact efficiency margins.[1][2][3]

Looking ahead, OpenAI plans to begin deploying the Jalapeño processor in small volumes within its own data centers by the end of 2026, with a broader production ramp scheduled throughout 2027. A second-generation version of the chip is already in fabrication, promising an additional 25 percent improvement in performance per watt. Up to 2,048 Jalapeño chips can be connected across a 16-rack scale-up domain, supporting OpenAI's ambitious goal of building massive, gigawatt-scale data centers that can handle the next generation of artificial intelligence demand.[4][7]

The custom processor is designed to deliver faster response times for highly interactive AI tasks.

The introduction of the Jalapeño processor signals a broader structural shift in the artificial intelligence infrastructure market. As the economics of inference become central to long-term profitability, major AI laboratories are increasingly seeking autonomy over their hardware stacks to control costs. While Nvidia's data center revenue continues to soar and its training hardware remains indispensable, the success of custom silicon like Jalapeño suggests that the future of AI inference may be driven by highly specialized, in-house processors designed to make computing power more abundant and affordable.[6][7]

Key points

  • OpenAI unveiled benchmark results for its first custom AI inference chip, Jalapeño, co-developed with Broadcom.
  • The processor reportedly delivers up to 1.9 times more work per watt and significantly lower latency than Nvidia's GB300 systems.
  • Jalapeño is designed exclusively for inference, allowing OpenAI to reduce the massive electricity costs of running large language models.
  • Despite the custom silicon breakthrough, OpenAI will continue to rely on Nvidia hardware for the computationally intensive task of training models.

Viewpoints in depth

OpenAI's Hardware Strategy

OpenAI views custom silicon as essential for scaling its operations and reducing the cost of serving AI models.

For OpenAI, the Jalapeño chip represents a critical step toward vertical integration. By designing the processor specifically for the inference workloads generated by its own models, the company can eliminate the inefficiencies of general-purpose hardware. Hardware chief Richard Ho emphasized that the chip's architecture minimizes data movement, a major source of power consumption. This full-stack approach—where the models themselves assist in designing the chips that will run them—allows OpenAI to optimize performance and potentially lower the cost of API access for developers and end users.

The Financial Market Perspective

Investors are weighing the chip's efficiency gains against OpenAI's ongoing dependence on Nvidia.

Financial analysts point out a paradox in OpenAI's announcement: while Jalapeño beat Nvidia's GB300 in inference benchmarks, OpenAI simultaneously reaffirmed its commitment to buying Nvidia hardware. Nvidia remains the undisputed leader in AI training, a workload Jalapeño is not designed to handle. Furthermore, Nvidia's massive supply commitments and recent agreement to backstop financing for OpenAI's data centers underscore a deeply intertwined relationship. Analysts suggest that while Jalapeño may reduce OpenAI's inference costs, it is unlikely to meaningfully displace Nvidia's dominance in the broader AI infrastructure market in the near term.

Semiconductor Industry Analysts

Hardware experts acknowledge the engineering achievement but highlight the missing comparison to Nvidia's newest chips.

Semiconductor researchers, including those at SemiAnalysis who facilitated the public InferenceX benchmarks, praise Jalapeño as a genuine engineering feat, particularly its use of HBM4 memory to achieve 15.4 terabytes per second of bandwidth. However, they caution that the benchmark comparisons are not entirely apples-to-apples. Jalapeño was tested against Nvidia's GB200 and GB300 systems, but not against the newer Vera Rubin architecture, which also utilizes HBM4 memory and is already shipping to customers. Industry observers note that by the time Jalapeño reaches volume production in 2027, Nvidia's hardware will have advanced further, potentially narrowing the current efficiency gap.

Why this matters

Inference—the process of running an already-trained AI model to answer user prompts—consumes vast amounts of electricity and is a primary driver of data center costs. By designing a custom chip that handles this specific workload more efficiently than general-purpose hardware, OpenAI aims to drastically lower the cost of operating its models, potentially translating to cheaper API access and faster response times for end users.

How we got here

  1. October 2025

    OpenAI and Broadcom announce a partnership to co-develop custom AI accelerators.

  2. November 2025

    The first Jalapeño design reaches manufacturing tape-out after a nine-month development cycle.

  3. August 2026

    OpenAI publishes the first public benchmark results for the Jalapeño chip at the Hot Chips conference.

  4. Late 2026

    OpenAI plans to begin deploying the first Jalapeño chips in small volumes within its data centers.

Sources

Source coverage

7 outlets

3 viewpoints surfaced

OpenAI Hardware Team 35%Financial Analysts 35%Semiconductor Industry Observers 30%
  1. [1]Tom's HardwareSemiconductor Industry Observers

    OpenAI says its Jalapeño chip beats Nvidia's GB300 in first published benchmarks

    Read on Tom's Hardware
  2. [2]24/7 Wall St.Financial Analysts

    OpenAI's Custom Chip Embarrasses Nvidia, While Company Vows to Keep Buying From It

    Read on 24/7 Wall St.
  3. [3]LiveMintSemiconductor Industry Observers

    OpenAI Jalapeño chip beats Nvidia Blackwell

    Read on LiveMint
  4. [4]Proactive InvestorsFinancial Analysts

    OpenAI's new AI chip outperforms Nvidia's GB300 in efficiency tests, company says

    Read on Proactive Investors
  5. [5]OpenAIOpenAI Hardware Team

    Jalapeño leads across GPT‑OSS operating points

    Read on OpenAI
  6. [6]Times of AISemiconductor Industry Observers

    OpenAI Says Its Jalapeño Chip Beats Nvidia's GB300 on Efficiency

    Read on Times of AI
  7. [7]ChannelchekSemiconductor Industry Observers

    OpenAI's New AI Chip Outperforms Nvidia's GB300 in Two Key Benchmarks

    Read on Channelchek

Comments

Stay informed

Every angle. Every day.

Get ai stories with full source coverage and perspective breakdowns delivered to your inbox.