Marcus Chen, YuSMP Group
Marcus Chen Staff Engineer, Backend & Cloud, YuSMP Group · Cloud infrastructure and AI deployment for US/EU products
Modern data center corridor with rows of GPU server racks illuminated by white LED strips, representing enterprise AI inference infrastructure

The short answer

IBM and Together AI announced a $240 million multi-year deal on August 11, 2026 to deploy a large-scale open-source AI inference cluster on IBM Cloud powered by Nvidia Blackwell-generation GPUs. The cluster — built on Nvidia HGX B300 systems — will be the first of its kind on IBM Cloud and is targeted for commercial availability in Q1 2027.

Together AI currently serves approximately 400 trillion tokens per month across more than 1 million developers. The deal is a direct response to enterprise demand for open-model alternatives to proprietary AI APIs: lower costs, transparent weights, and infrastructure not locked to a single vendor.

What the deal actually involves

IBM and Together AI signed what they describe as a multi-year infrastructure agreement, with IBM Cloud hosting Together AI's inference platform on approximately 2,000 Nvidia Blackwell 300 chips. IBM positioned the deal as the first dedicated large-scale cluster built for inference on IBM Cloud using HGX B300 systems — a distinction that matters because inference workloads have different resource profiles than training: they require fast, low-latency throughput at scale rather than sustained high-utilization training runs.

Together AI's CEO Vipul Ved Prakash framed the business case directly: "Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale." IBM Cloud General Manager Alan Peacock described the cluster as supporting "agentic AI at scale to drive real business outcomes."

Together AI raised an $800 million Series C at an $8.3 billion valuation as of July 2026, which gives the company the runway to execute multi-year infrastructure commitments of this scale. The IBM deal adds a hyperscaler-adjacent partner and brings IBM's enterprise sales relationships to Together AI's technical platform.

Why open-source inference is winning enterprise interest

The strategic argument for open-source inference has shifted from philosophical to practical over the past year. Three factors are driving the change inside enterprise engineering organizations.

Cost predictability. Proprietary AI API pricing is controlled by the vendor and can change — often upward — as products mature. An inference cluster built on open-source models lets engineering teams negotiate infrastructure contracts with fixed pricing rather than accepting per-token rates set by an API provider. At 400 trillion tokens per month across Together AI's platform, even a fraction of a cent in cost difference per thousand tokens represents material budget impact.

Model transparency and security review. Open-source model weights can be audited, inspected, and — for organizations with the infrastructure — hosted entirely within a controlled environment. Proprietary models are opaque by definition: you cannot verify what data influenced the model, how it handles sensitive inputs, or whether it has been modified after initial release. For FinTech and HealthTech teams handling regulated data, this transparency gap is a compliance risk, not just a philosophical preference.

Vendor diversification. European enterprises in particular face growing pressure to reduce dependence on a small number of US hyperscalers for AI capability. The IBM-Together AI cluster offers an additional procurement path — one that IBM can position within its existing enterprise relationships and data-residency commitments.

Technical specifications of the cluster

The cluster is built on Nvidia HGX B300 systems — the Blackwell-generation GPU configuration optimized for high-throughput inference and distributed training. Nvidia's Dion Harris described these systems as part of an "AI factories" model: "AI factories are becoming essential enterprise infrastructure — like electricity and telecommunications." Nvidia claims 30x improvement in AI factory output versus prior-generation hardware configurations.

Network fabric is Nvidia Spectrum-X Ethernet, which is optimized for AI workloads where traditional datacenter Ethernet fabrics show congestion under the bursty all-to-all communication patterns of large model inference. The U.S.-based deployment means the cluster is initially scoped for the North American market.

Together AI's platform supports training and reinforcement learning in addition to inference, giving enterprises building fine-tuned models a single-vendor path for the full model lifecycle — not just serving. This is relevant for teams doing cloud-native AI deployments where training a domain-specific variant of an open model and then serving it at scale requires cohesive infrastructure.

What it means for US & EU software teams

The IBM-Together AI deal is a signal, not a solution. It does not immediately change the cost or capability of open-source AI inference for most teams — Q1 2027 availability means there is runway before this cluster is an operational option. But it shifts the competitive landscape in ways that matter for teams making AI infrastructure decisions today.

For US teams: The deal validates open-source inference as a viable enterprise path alongside — not instead of — proprietary APIs. Teams that have been evaluating Together AI's platform on shared infrastructure will, by Q1 2027, have access to dedicated Nvidia Blackwell capacity backed by IBM Cloud SLAs. This is relevant for startups and mid-market engineering organizations that currently accept shared GPU resources because they cannot afford dedicated hardware contracts.

For EU teams: The cluster is U.S.-based, which does not resolve GDPR data residency requirements directly. However, the open-source model architecture means EU teams can run equivalent configurations on IBM Cloud EU regions — or on other compliant infrastructure — using the same model weights. The deal also reduces the AI market concentration risk that EU regulators have flagged: a credible IBM Cloud inference path means enterprises have a genuine alternative to Azure OpenAI, AWS Bedrock, and Google Vertex AI for AI workloads.

The practical question for engineering leadership: if your team's AI API spend is growing faster than your engineering budget, and a meaningful portion of that spend is on inference rather than training, open-source alternatives running on enterprise-grade infrastructure deserve a structured evaluation now — before the budget conversation becomes urgent.

Frequently asked questions

What is the IBM and Together AI $240M deal?

IBM and Together AI signed a multi-year, $240 million agreement announced on August 11, 2026, to build a large-scale AI inference cluster on IBM Cloud. The cluster uses Nvidia HGX B300 systems — Blackwell-architecture GPUs — paired with Nvidia Spectrum-X Ethernet networking. It is the first dedicated, large-scale inference cluster built on IBM Cloud using this hardware generation. Commercial availability is targeted for Q1 2027.

Which open-source AI models will run on the cluster?

Together AI's platform supports a broad catalog of open-source models. Current examples include DeepSeek, MiniMax, and Kimi models. Together AI serves approximately 400 trillion tokens per month across more than 1 million developers globally, with the new IBM Cloud cluster intended to expand capacity for production inference and reinforcement learning workloads.

When will the IBM-Together AI cluster be available?

IBM said the cluster is expected to reach commercial availability in Q1 2027. The infrastructure build is already underway, with approximately 2,000 Nvidia Blackwell 300 chips deployed in a U.S.-based facility.

How does open-source inference compare in cost to proprietary AI APIs?

Open-source inference typically offers lower per-token costs at scale compared to proprietary APIs, though the gap depends on model size and use case. The strategic advantage is predictable pricing and the ability to negotiate infrastructure contracts rather than being subject to vendor-controlled API pricing. Together AI's CEO noted that enterprises "want the performance of the best frontier models without the closed-model price tag."

What does this deal mean for EU teams with data sovereignty requirements?

The IBM-Together AI cluster is U.S.-based, so it does not directly resolve EU data residency requirements under GDPR. However, open-source model weights can be deployed on compliant EU infrastructure — IBM Cloud EU regions or sovereign cloud alternatives — using the same models. The deal also reduces vendor lock-in to a single American hyperscaler, which is relevant for EU enterprises under the European Data Act and forthcoming Cloud Switching Regulation.

Sources

IBM Newsroom — IBM and Together AI Sign Multi-Year Agreement to Scale Open-Source AI Inference, August 11, 2026
BNN Bloomberg — IBM, Together AI ink $240 million deal for Nvidia-powered AI inference cluster, August 11, 2026
The Next Web — IBM bets $240m on cheap, open-source inference to take on the hyperscalers, August 2026