Nvidia launches Groq 3 LPX inference accelerator in full production

August 24, 2026 11:00 AM EDT

Nvidia (NASDAQ: NVDA) announced that its Groq 3 LPX AI inference accelerator is now in full production, according to a company press release issued at the Hot Chips conference.

The Groq 3 LPX is designed as an extension of the Nvidia Vera Rubin NVL72 platform, with the stated purpose of increasing token generation rates for agentic AI workloads. In benchmarking conducted by Artificial Analysis, the accelerator recorded 3,400 output tokens per second running Gemma 4 31B, an open-source model, with a 100,000-token context window. Nvidia described this as the fastest performance recorded for that model.

Nvidia claims the Groq 3 LPX provides four times faster responsiveness for latency-sensitive workloads compared to what it describes as the nearest alternative platform.

"Inference is the growth engine of AI," said Jensen Huang, founder and CEO of Nvidia. "Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation."

Nebius, an AI cloud provider, is identified as the first cloud to adopt Groq 3 LPX, with plans to deploy it through its Nebius Token Factory inference platform.

"As the first AI cloud bringing it to production via Nebius Token Factory, we're making sure every step of an agent's loop feels instant," said Danila Shtan, chief technology officer of Nebius.

Groq, a purpose-built AI inference cloud company, is also named as among the platform's earliest planned adopters following Nebius.

The Groq 3 LPX works alongside Nvidia BlueField-4 DPUs, Vera CPU racks, Vera BlueField-4 STX storage, and Nvidia Spectrum-6 SPX Ethernet within rack-scale configurations.



Serious News for Serious Traders! Try StreetInsider.com Premium Free!

You May Also Be Interested In





Related Categories

Corporate News

Related Entities

Maynard Um, Mark Zuckerberg, ARK