INFRA Signal 116
Nvidia says Groq 3 LPX racks hit 3,400 tokens per second in Artificial Analysis benchmark, enter full production
Nvidia reports its Groq 3 LPX inference racks achieved 3,400 tokens per second on an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence, and the hardware has entered full production with Nebius as the first customer.
This is the first public benchmark result from Nvidia's $20 billion Groq licensing deal, and a tweet from an observer claims the throughput is 4x the next public endpoint. For teams building inference infrastructure, the production availability of dedicated LPU-based racks introduces an alternative to GPU-based serving, though the material provides no pricing or availability details beyond Nebius as an early adopter.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Groq 3 LPX racks delivered 3,400 tokens per second running Gemma 4 31B with a 100,000-token input sequence in an Artificial Analysis benchmark.
The hardware has entered full production, with Nebius Token Factory named as the first AI cloud to adopt it.
Nvidia's $20 billion Groq licensing deal underpins the product, and Groq itself will deploy Groq 3 LPX alongside Nvidia Vera Rubin NVL72.
THE CLUSTER
↗