**IBM and Together AI Partner to Deploy Major NVIDIA AI Inference Cluster on IBM Cloud**
IBM is set to deploy a large-scale AI computing cluster on IBM Cloud in a multiyear agreement valued at $240 million with Together AI. Expected to become available in the first quarter of 2027, the cluster will be built using NVIDIA HGX B300 systems interconnected via NVIDIA Spectrum-X Ethernet networking. Together AI will leverage this infrastructure to run inference workloads for open-source AI models.
The initial deployment will feature approximately 2,000 NVIDIA Blackwell 300 GPUs and will be located in the United States. According to Together AI’s chief revenue officer, the capacity is expected to be fully committed within two to three months of its launch. This deployment marks IBM Cloud’s first dedicated large-scale inference cluster centered on HGX B300 systems. NVIDIA claims that the HGX B300 and Spectrum-X configuration is designed to deliver 30 times more AI factory output than previous generations.
The cluster is primarily intended for inference workloads, where trained AI models process requests and generate outputs. Inference has emerged as a major driver of demand for computing capacity, prompting cloud providers and chipmakers to expand their AI infrastructure offerings. Together AI provides infrastructure and software for AI inference, training, fine-tuning, and agent-based workloads. The company recently raised $800 million in a Series C funding round, valuing it at $8.3 billion.
In addition to the IBM agreement, Together AI has secured commitments for more than 500 MW of compute capacity, which will be financed independently by new investors. The company’s inference service currently processes more than 400 trillion tokens each month, with monthly token volume growing from 30 billion to 400 trillion in just nine months.
To better serve customers, Together AI launched a Provisioned Throughput service in July, allowing users to reserve a defined rate of model processing measured in tokens per minute. Customers purchase Provisioned Throughput Units (PTUs), with each unit representing a fixed slice of guaranteed throughput for a selected model or model family. Together AI manages the underlying infrastructure while providing reserved token-processing capacity through its API.
The IBM agreement adds another dedicated pool of computing resources to this infrastructure. Together AI selected IBM and NVIDIA based on their GPU capacity and respective infrastructure roadmaps.
Enterprises increasingly seek the performance of top-tier AI models without the cost of proprietary solutions, and this partnership aims to deliver exactly that. “Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale,” said Vipul Ved Prakash, CEO of Together AI. “Working alongside IBM with NVIDIA gives us that foundation. This cluster lets us bring production-grade inference to more companies, faster, and it’s a big step in our push to make open-source AI the obvious choice for enterprises.”
IBM’s expansion of its NVIDIA infrastructure partnership is not entirely new. In March 2026, IBM announced plans to make NVIDIA Blackwell Ultra GPUs available through IBM Cloud, supporting large-scale model training, high-throughput inference, and AI reasoning. The companies are also collaborating on high-performance storage and data center solutions, with IBM Storage Scale System 6000 selected to provide 10 PB of storage for NVIDIA’s GPU-based analytics systems. The partnership also explores integrations with IBM Sovereign Core and Nemotron models for GPU-intensive workloads requiring regional data residency.
This latest deployment with Together AI continues IBM’s focus on inference workloads using dedicated infrastructure. Meanwhile, other major cloud providers are also investing heavily in NVIDIA technology. AWS, for example, has agreed to purchase one million NVIDIA GPUs by the end of 2027 and plans to deploy NVIDIA ConnectX and Spectrum-X networking equipment in its data centers.
—
### FAQ
**What is the IBM and Together AI partnership about?**
IBM and Together AI have entered a multiyear, $240 million agreement to deploy a large NVIDIA-based AI inference cluster on IBM Cloud. The cluster, expected in Q1 2027, will use NVIDIA HGX B300 systems and Spectrum-X networking to support open-source AI model inference.
**What hardware will the cluster use?**
The cluster will initially include about 2,000 NVIDIA Blackwell 300 GPUs and will be located in the United States. It will be IBM Cloud’s first dedicated large-scale inference cluster built around NVIDIA HGX B300 systems.
**What is inference, and why does it matter?**
Inference is the process where trained AI models generate outputs in response to input requests. It has become one of the largest drivers of demand for computing capacity, prompting cloud providers and chipmakers to expand AI infrastructure dedicated to inference workloads.
**What is Together AI’s Provisioned Throughput service?**
Launched in July, this service allows customers to reserve a defined rate of model processing measured in tokens per minute. Customers purchase Provisioned Throughput Units (PTUs), each representing a guaranteed slice of throughput for a selected model or model family, while Together AI manages the underlying infrastructure.
**How does this partnership fit into IBM’s relationship with NVIDIA?**
This deployment extends IBM’s existing collaboration with NVIDIA. In March 2026, IBM announced plans to make NVIDIA Blackwell Ultra GPUs available through IBM Cloud. The companies are also working together on high-performance storage and integrating IBM Sovereign Core with NVIDIA infrastructure for regulated workloads.
**What are the broader industry trends reflected in this partnership?**
The agreement highlights the growing demand for dedicated AI inference capacity, with enterprises racing to adopt agentic AI at scale. It also reflects the trend of cloud providers combining internally developed infrastructure with NVIDIA technology to meet rising AI compute needs.
—
### Conclusion
The partnership between IBM, Together AI, and NVIDIA marks a significant step in bringing large-scale AI inference capabilities to the cloud. By deploying a dedicated cluster centered on NVIDIA HGX B300 systems, IBM and Together AI aim to deliver production-grade, open-source AI inference at scale and speed. This development not only strengthens IBM’s position in the AI infrastructure market but also underscores the growing importance of flexible, high-performance computing resources in supporting the rapid adoption of AI technologies across industries. As demand for AI capacity continues to surge, collaborations like this will play a crucial role in shaping the future of enterprise AI.



