Nvidia Groq 3 LPU chips during the Nvidia GTC conference on March 18, 2026 in San Jose, California.
David Paul Morris | Bloomberg | Getty Images
Nvidia announced Monday that its Groq 3 LPX rack is in full production, marking the commercialization of the technology through the company’s largest acquisition in its history.
Groq racks will be introduced on neocloud alongside Vera central processors and Rubin graphics processors. NeviusNvidia senior director Dion Harris told reporters it will be online later this year.
Nvidia’s race to manufacture and deliver Groq chips to customers highlights the growing importance of low-latency inference, which is needed to make AI agents feel responsive without long delays for users, especially when coding. Nvidia says cloud companies can charge additional fees for these types of tokens.
“It unlocks the ability for those who are offering the token to be able to offer a premium tier of service to users and customers who actually demand the most latency-sensitive service contracts,” Harris said on the conference call.
In December, Nvidia acquired assets from chip startup Groq for $20 billion, its largest acquisition ever.
The Groq architecture includes 500 megabytes of high-speed SRAM on the chip’s die itself to alleviate memory-related bottlenecks. The Groq chip is made by Samsung. taiwan semiconductor manufacturing Manufactures GPUs for Nvidia.
Nvidia packages 256 individual Groq 3 chips into an LPX rack. Citing Artificial Analysis benchmarks, Nvidia said its Groq 3 LPX rack can deliver 3,400 tokens per second.
It’s a competitive place. small GPU manufacturers advanced micro device announced earlier this year that it would integrate the chip with its rack-scale systems. cerebrumrecently published, focuses on low-latency inference. OpenAI’s newly announced Ultrafast mode currently promises to issue 750 tokens per second and is “powered by Cerebras.”
Low-latency chips do not replace GPUs, the workhorse of AI chips, which can perform not only inference but also training, and are flexible enough to adapt to new technologies and models. Low-latency chips like Groq primarily focus on a part of the service delivery model called the “decode” phase.
“This is not about replacing GPUs,” Harris said. “The key is to use the right processor at the right price for the right part of the workload.”
Nvidia is currently ramping up shipments of its Vera Rubin systems, which began production earlier this year. At the Vera Rubin and Groq 3 LPX launch event in March, Nvidia CEO Jensen Huang predicted that the cumulative sales of current-generation Blackwell chips and new Vera Rubin systems would reach $1 trillion by 2027.
Huang said at the time that the company would allocate a quarter of its data center space to Groq chips for application coding.
“The rest of my data center is all 100% Vera Rubin,” Huang says.
Nvidia will announce its financial results on Wednesday.
Watch: Inside Nvidia’s Vera Rubin AI system

