ELSEIF
Your brief EB
335 stories from 95 feeds 232 clusters Refreshed 5 minutes ago next pull 06:21

TECH Signal 296

Together AI secures $240M IBM Cloud deal to deploy Nvidia HGX B300 GPUs for inference in Q1 2027

Together AI will use IBM Cloud to host a large-scale deployment of Nvidia’s HGX B300 GPUs for AI inference workloads starting early 2027.

WHY IT MATTERS

This deal highlights the ongoing scramble for AI compute capacity, even if it means relying on last-gen hardware. For engineers, it signals that cloud providers are willing to broker large-scale GPU deployments to meet demand, regardless of vendor competition. The choice of B300s over newer Nvidia systems suggests cost and availability may outweigh raw performance for some workloads.

Written by elseif from the cluster below · every claim links back to a source

The three things worth knowing

01

Together AI will deploy Nvidia’s HGX B300 GPUs on IBM Cloud for AI inference, funded by a $240M deal.

02

The B300 platform, while not Nvidia’s latest, is designed for conventional air-cooled datacenters and inference workloads.

03

The deployment is scheduled for Q1 2027, reflecting the lead time required for large-scale GPU infrastructure rollouts.

THE READ

What the cluster adds up to.

ORIGINAL ANALYSIS

Together AI’s $240M deal with IBM Cloud underscores the intense demand for AI compute capacity, even when it involves partnering with a competitor. The company, which primarily rents GPU resources from cloud providers, is expanding its infrastructure to include a large cluster of Nvidia’s HGX B300 GPUs. This move is less about cutting-edge hardware and more about securing the necessary capacity to scale its open-weights inference platform. For engineers, this deal is a reminder that availability and cost-efficiency often trump the latest specifications in production deployments.

The choice of Nvidia’s HGX B300 platform is notable. Unlike Nvidia’s higher-end systems, which require specialized cooling and power infrastructure, the B300 is designed for traditional air-cooled datacenters. Each 14-15 kW box contains eight GPUs interconnected via NVLink and Spectrum-X Ethernet, making it a practical choice for inference workloads. IBM’s announcement that this is its first large-scale B300 deployment for inference suggests that the platform is gaining traction for cost-sensitive, high-volume use cases. However, the trade-off is performance, engineers should expect lower throughput compared to Nvidia’s latest offerings.

The timeline for this deployment, Q1 2027, highlights the long lead times involved in procuring and deploying large-scale GPU infrastructure. This delay is likely due to supply chain constraints, datacenter capacity limitations, and the logistical challenges of integrating new hardware. For engineers planning AI workloads, this serves as a cautionary note: securing compute resources may require advance commitments, and last-minute scaling could be difficult. The deal also reflects a broader trend where cloud providers act as intermediaries, brokering GPU capacity to meet demand even if it means working with competitors.

Together AI’s flexibility in hardware selection, deploying on both Nvidia GPUs and SambaNova’s heterogeneous compute platform, demonstrates that the company prioritizes price-performance over vendor lock-in. This approach allows it to optimize costs for inference workloads, which are typically less demanding than training. However, engineers should note that this flexibility comes with trade-offs, such as potential compatibility issues or reduced optimization for specific hardware. The deal’s focus on inference also suggests that training workloads may still require more specialized or higher-end infrastructure.

Written by elseif from the cluster below · checked for specifics the sources never contained

THE CLUSTER

Same story, 1 feed.

ORDERED BY FIRST SEEN
www.theregister.com - Articles Together AI embraces the competition with $240M IBM Cloud deal Open ↗