Together AI says it is expanding its enterprise inference capacity with a dedicated cluster of NVIDIA B300 GPUs running on IBM Cloud and connected with NVIDIA Spectrum-X Ethernet. Together AI will operate the inference layer, IBM will provide the cloud infrastructure and NVIDIA supplies the accelerator and networking technology.

The October update builds on an agreement IBM announced in August. IBM described a multi-year $240 million arrangement with Together AI for a large-scale inference cluster using NVIDIA HGX B300 systems and Spectrum-X networking, with expected availability in the first quarter of 2027. Together AI's new post frames the deployment as the first dedicated large-scale inference cluster of its kind on IBM Cloud and says it will be used to serve open models for enterprise workloads.

The architecture reflects the growing specialization of AI infrastructure. Training clusters are optimized for building models, while high-volume inference requires sustained token throughput, low latency, predictable reliability and efficient networking. Together AI's role is to provide the model-serving layer on top of IBM's infrastructure, while NVIDIA's B300 systems and Spectrum-X handle accelerated compute and network fabric.

Together AI says it already serves hundreds of trillions of tokens per month to more than one million developers. Those usage figures are vendor-reported and were not independently validated in this run. IBM's earlier announcement does, however, independently establish the commercial agreement from the perspective of another party to the deal, including the planned B300 deployment, the $240 million value and the Q1 2027 expected availability.

The development is important because demand for inference is becoming an infrastructure market of its own. Enterprises adopting open models need capacity that can be scaled without giving up deployment control, while inference providers need access to current accelerators and high-bandwidth networking. The partnership joins a model-serving specialist, a major cloud operator and the dominant supplier of AI accelerators. The remaining question is execution: the cluster's production economics, delivered performance and actual availability will only be clear once the planned infrastructure is online and independently observable.

References