Skip to main content
AI InfrastructureExplainerAug 29, 2026, 9:30 AM· 4 min read· in technology

IBM and Together AI Sign $240 Million Deal to Build Massive Nvidia B300 Inference Cluster

IBM and neocloud provider Together AI are partnering to deploy a large-scale Nvidia AI inference cluster on IBM Cloud. The $240 million investment aims to drive down the cost of running open-source AI models for enterprise customers by early 2027.

By Wei Zhang

Enterprise Cloud Providers 35%Neocloud & Open-Source Advocates 35%Hardware & Market Analysts 30%
Enterprise Cloud Providers
Argue that AI factories are becoming essential utilities, requiring the stability and security of established enterprise clouds to move from pilot to production.
Neocloud & Open-Source Advocates
Emphasize that open, modular stacks are the future of AI, and that driving down token costs is necessary to make agentic workflows accessible to all businesses.
Hardware & Market Analysts
Point out that while the demand is real, the long lead time to 2027 introduces pricing and obsolescence risks, turning AI infrastructure into a highly competitive arms race.

Key terms

Inference
The process where a trained AI model generates responses, predictions, or decisions based on new data.
Neocloud
A cloud computing provider that specializes specifically in AI workloads and GPU infrastructure, rather than traditional web hosting.
Token Economics
The financial cost associated with processing or generating a single unit of data (a token) in an AI model.
Agentic AI
Artificial intelligence systems designed to autonomously plan and execute multi-step tasks to achieve a specific goal.

Key points

  1. IBM and Together AI have signed a $240 million multi-year agreement to build an AI inference cluster on IBM Cloud.
  2. The deployment will utilize Nvidia's next-generation HGX B300 systems and Spectrum-X Ethernet networking.
  3. Together AI will use the infrastructure to serve open-source AI models to enterprise customers at scale.
  4. The cluster is currently in the planning and buildout phase, with availability targeted for the first quarter of 2027.
  5. The partnership aims to improve token economics, making it cheaper for businesses to run complex, real-time AI workloads.

For enterprises attempting to integrate artificial intelligence into their daily operations, the primary bottleneck is no longer the intelligence of the models themselves. Instead, it is the sheer cost and availability of the computing power required to serve them. As open-source models become highly capable alternatives to proprietary systems, the infrastructure needed to run them at scale is transitioning into a critical business utility.

To address this growing demand, IBM and neocloud provider Together AI have signed a $240 million multi-year agreement to construct a massive AI inference cluster on IBM Cloud. The partnership aims to provide enterprise customers with the dedicated computing muscle necessary to run open-source models efficiently and reliably.[1][2][4]

The planned deployment will be powered by Nvidia's next-generation HGX B300 systems, linked together by Nvidia Spectrum-X Ethernet networking. According to IBM, this represents its first dedicated, large-scale inference cluster utilizing this specific hardware combination on its cloud platform.[2][4][5]

However, while the dollar figure and hardware specifications are substantial, this announcement is fundamentally a capacity bet for the future rather than a live service today. The cluster is currently in the planning stages and is not expected to become available to customers until the first quarter of 2027.[5][6]

The $240 million deployment is targeted to go live in the first quarter of 2027.

To understand the mechanics of the deal, it is necessary to look at Together AI's role as a "neocloud." Unlike traditional hyperscalers that offer a vast array of general computing services, neoclouds are specialized providers that rent out AI-specific computing capacity.[3]

Instead of businesses purchasing their own costly graphics processing units (GPUs) and managing the complex networking required to link them, they use platforms like Together AI to run, customize, and fine-tune open-source models on demand.[1][3]

The scale of the demand for this specialized infrastructure is staggering. Together AI, which recently raised $800 million at an $8.3 billion valuation, reports that its platform is already processing roughly 400 trillion tokens per month.[1][5]

The scale of the demand for this specialized infrastructure is staggering.

For Together AI, partnering with a legacy technology giant provides crucial enterprise credibility. While startups and developers are comfortable using emerging neoclouds, Fortune 500 companies often require the stability, security, and proactive service guarantees associated with established providers like IBM.[1]

Conversely, the agreement offers IBM a lucrative entry point into the fast-growing neocloud market without assuming the immense capital risks of becoming a pure-play AI data center operator. IBM secures a major tenant for its cloud infrastructure while bolstering its generative AI service backlog.[3]

Neoclouds specialize in renting out AI-specific computing capacity rather than general cloud services.

The underlying engine of this infrastructure buildout is Nvidia's B300 architecture. Nvidia claims that the B300 platform is designed to deliver up to 30 times the "AI factory output" of previous generations.[2][4]

Yet, that 30x figure requires careful framing. It is a platform-level design claim from the manufacturer, not a measured performance result from this specific, yet-to-be-built cluster. The actual efficiency gains will depend heavily on how the networking and software layers are optimized once the hardware is installed.[5]

The ultimate objective of the partnership is to improve "token economics"—the financial cost of generating each word or piece of data. As inference workloads grow and businesses experiment with autonomous, multi-step agentic AI, reducing the cost per token is essential for making these tools financially viable.[1][2][4]

Enterprises are increasingly looking to open-source AI models to maintain control over their data.

This infrastructure is specifically targeted at serving open-source models, which are gaining traction among enterprises seeking more control over their data and customization options. By providing enterprise-grade hosting for these models, the cluster aims to give companies a robust alternative to being locked into proprietary ecosystems.[1][4][5]

The long lead time, however, introduces inherent market risks. By the time the cluster goes live in early 2027, the B300 hardware will be competing with even newer architectures, and hyperscalers will have vastly expanded their own capacities.[6]

Ultimately, the deal underscores a broader shift in the technology sector. AI computing is rapidly transitioning from an experimental luxury into a fundamental utility, and the race to build the factories that will power it is only just beginning.[2][4]

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Enterprise Cloud Providers 35%Neocloud & Open-Source Advocates 35%Hardware & Market Analysts 30%
  1. [1]Fierce NetworkNeocloud & Open-Source Advocates

    Together AI taps IBM Cloud in a $240M compute deal to scale open-source AI inferencing on Nvidia GPUs

    Read on Fierce Network
  2. [2]Pulse 2.0Enterprise Cloud Providers

    IBM And Together AI Sign $240 Million Multi-Year Deal For Large-Scale NVIDIA B300 AI Inference Cluster

    Read on Pulse 2.0
  3. [3]MarketWiseHardware & Market Analysts

    IBM's $240M Together AI Neocloud Deal – What It Means for AI Stocks

    Read on MarketWise
  4. [4]IBM NewsroomEnterprise Cloud Providers

    IBM and Together AI Sign Multi-Year Agreement to Scale Open-Source AI Inference with NVIDIA AI Infrastructure on IBM Cloud

    Read on IBM Newsroom
  5. [5]Superpower DailyNeocloud & Open-Source Advocates

    Together AI is lining up a dedicated, large-scale IBM Cloud cluster to serve open-source models to enterprise customers

    Read on Superpower Daily
  6. [6]Semiconductor NewsHardware & Market Analysts

    IBM Puts $240M Behind NVIDIA B300 for AI Inference

    Read on Semiconductor News

Comments

Stay informed

Every angle. Every day.

Get technology stories with full source coverage and perspective breakdowns delivered to your inbox.