INFRA Signal 183
SpaceXAI to add 660,000 GB300 GPUs this year, reaching 1.44 million total
Elon Musk announced that SpaceXAI will deploy three tranches of 220,000 Nvidia GB300 GPUs by late December, bringing its total operational count to 1.44 million units.
This expansion resolves the hardware bottleneck for training Grok by consolidating compute on Blackwell architecture, though it introduces significant power infrastructure challenges. The shift away from mixed-architecture clusters to homogeneous Blackwell sites eliminates training inefficiencies but relies on controversial on-site power generation.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Colossus 2 will host 1.1 million GB300 GPUs, while Colossus 1 retains a mixed Hopper and Blackwell configuration.
SpaceXAI is using unpermitted gas turbines to power the expansion, with a lawsuit filed by local residents over the noise and emissions.
The company plans to grow data center capacity sevenfold by 2027, targeting 50 million H100-equivalent GPUs by 2030.
THE READ
What the cluster adds up to.
SpaceXAI is transitioning its compute infrastructure from a mixed-architecture environment to a homogeneous Blackwell cluster. Colossus 1, which combines H100, H200, and GB200 GPUs, is described as inefficient for training Grok due to architectural bottlenecks. Consequently, the company has rented Colossus 1 to Anthropic for inference workloads, freeing up the newer Colossus 2 site for dedicated training tasks.
The hardware expansion is substantial, adding 660,000 GB300 GPUs in three tranches of 220,000 units each. This brings the total operational count to 1.44 million GPUs, with 1.1 million of those being GB300 units at Colossus 2. This scale allows SpaceXAI to match the compute density of larger, older competitors like OpenAI and Anthropic, despite being only three years old.
The primary constraint for this build-out is not GPU availability but power generation. SpaceXAI has installed its own gas turbines to meet the energy demands of the data center, a move that has triggered legal action from local residents. The company has promised to remove the turbines within one year, contingent on the completion of its own 1.2-GW power plant.
Future plans indicate a continued aggressive expansion, with Musk stating the company aims for a sevenfold increase in capacity by 2027. Long-term targets include 50 million H100-equivalent GPUs by 2030, potentially involving orbital data centers. These projections highlight a strategy focused on vertical integration of both compute and energy infrastructure to sustain rapid scaling.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER
↗