IBM has committed to a multiyear US$240 million partnership with Together AI to deploy an NVIDIA HGX B300 GPU cluster on IBM Cloud, targeting high-performance AI training and open-model inference workloads for enterprise customers, the companies said this week.
A Blackwell-class cluster on IBM Cloud
The HGX B300 is NVIDIA's densest Blackwell-generation configuration, pairing eight B300 accelerators with fifth-generation NVLink and Spectrum-X networking. Together AI, best known for hosting Llama, DeepSeek, Qwen and other open-weight models at scale, will operate the cluster on behalf of IBM Cloud customers who need dedicated capacity for large-context inference, custom fine-tunes and multi-tenant serving.
Why enterprise buyers care about open models
The deal doubles down on IBM's bet that regulated buyers — banks, healthcare providers, defence contractors — will prefer open models they can inspect and self-host over black-box APIs. Together AI's inference stack lets those customers rent HGX B300 capacity without leaving IBM's compliance perimeter, competing head-on with Anthropic's US$45 billion Nscale deal in West Virginia and Anthropic's Theseus-backed infra push.
A busy month for AI-compute deals
Together AI's tie-up caps a run of headline compute deals: AWS committed to 2 million GPUs with NVIDIA's Vera CPU roadmap, AM Intelligence and Greenko announced a 9,000-Vera-Rubin AI factory, and SpaceX and xAI unveiled Vera-based Grok orbital plans. IBM's US$240M outlay is modest by comparison, but the choice of Together AI as delivery partner signals a strategic shift: hyperscalers are increasingly happy to source specialised inference stacks from focused independents rather than build every layer themselves.
Reporting based on coverage from HPCwire AIwire and Enterprise AI industry press.