IBM Cloud to Host $240M Together AI Inference Cluster

IBM and Together AI’s $240M NVIDIA infrastructure agreement shows how the enterprise partner ecosystem for open-model inference is expanding.

Written By
Eric Mboizi
Eric Mboizi
Aug 12, 2026
3 minute read
Channel Insider content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

Proprietary AI companies such as OpenAI have secured multibillion-dollar infrastructure agreements to meet growing demand. Together AI is now making a sizable infrastructure commitment of its own.

IBM and Together AI have signed a multi-year, $240 million agreement under which IBM plans to deploy a large cluster of NVIDIA HGX B300 systems on IBM Cloud. Expected to become available in the first quarter of 2027, the cluster will support Together AI’s open-model inference services.

The announcement comes about a month after Together AI raised $800 million in Series C funding from a group of investors that included NVIDIA and Salesforce Ventures.

Open models are ready for enterprise use

Two years ago, open models generally trailed proprietary alternatives, according to research from Epoch AI. Since then, newer open-weight models have narrowed the performance gap on some evaluations.

Moonshot AI’s Kimi K3 now competes with proprietary models on some software-engineering evaluations. In a Together AI analysis of the DeepSWE benchmark, which uses long-horizon tasks from open-source software repositories, Kimi K3 trailed Claude Fable 5 on first-attempt performance but surpassed it when the models were allowed multiple attempts. Because Together AI conducted the comparison, the results should be treated as company claims.

Open models are no longer limited to hobbyist projects. Together AI says its customers include Cursor, Cognition, and ElevenLabs, although the company does not publicly detail every customer’s workload or model mix.

In late 2025, Together AI said its platform ranked first in output-speed benchmarks for several leading open models, delivering up to twice the performance of some competing services.

The results demonstrated progress in open-model inference. However, speed alone does not establish enterprise maturity, which also depends on reliability, security, governance, and support.

Building for the future

In a video message shared by the CEO of Together AI, Vipul Ved Prakash, he says that the demand for AI compute is there but the main bottleneck at the moment is the economics of companies making such an investment. 

He argues closed source models extract “all the margin” out of AI investments and that open ecosystem plus the right infrastructure are the answer to this challenge.

Th NVIDIA HGX B300 hardware that the company plans to use is a strategic physical infrastructure choice for a number of reasons. First, the HGX B300 is a dual-use GPU, meaning that it can both be used for training AI models and inference tasks. This gives the Together AI team flexibility in future offerings since they are a full-stack AI platform and not just an inference service.

Advertisement

Second, the HGX B300 consists of 8 Blackwell GPUs, which is hardware that their inference engine has specifically been optimized to use.

Additionally, NVIDIA says that this GPU version is 30 times better than its earlier releases, offering faster speeds at a lower cost.

The enterprise takeaway

Recent improvements in open models like Kimi, Gwen and Deepseek have made it feasible for enterprises to consider model choices other than the proprietary ones.

The American telecom giant AT&T is one company that has recently made this shift, leading to savings of 80% to 90% in some applications. The company says that about 25% of its AI usage is now through open-weight or open-source models.

As Together AI makes the bet to build infrastructure for open model inference, enterprises are already showing the need for such offerings. Enterprises need to begin reviewing their LLM provider options in this new landscape. 

Read more: IBM’s recent revenue warning shows how enterprise customers are prioritizing scarce infrastructure.

Eric Mboizi

Eric Mboizi is a technology news writer covering software development, emerging technologies, and the evolving digital landscape for TechRepublic and eWeek. He holds a bachelor’s degree in software engineering from Makerere University and has more than five years of experience creating technical content for developers and technology professionals. In addition to his work as a journalist, Eric is an Ethereum developer with more than four years of experience in blockchain technology. His hands-on development background gives him a practical perspective on software engineering, decentralized technologies, and the real-world implications of new technology trends.

Channel Insider Logo

Channel Insider combines news and technology recommendations to keep channel partners, value-added resellers, IT solution providers, MSPs, and SaaS providers informed on the changing IT landscape. These resources provide product comparisons, in-depth analysis of vendors, and interviews with subject matter experts to provide vendors with critical information for their operations.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.