- In Artificial Analysis benchmarking, Groq 3 LPX showcased world-class speed for agentic coding and other latency-sensitive workloads.
- NVIDIA Groq 3 LPX extends the inference performance of NVIDIA Vera Rubin NVL72 systems by dramatically increasing token generation rates.
- Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX.
PALO ALTO, Calif., Aug. 24, 2026 (GLOBE NEWSWIRE) -- Hot Chips—NVIDIA today announced that NVIDIA Groq 3 LPX, the interactive AI inference accelerator, is now in full production. An extension of the NVIDIA Vera Rubin platform, Groq 3 LPX delivers a major boost in AI inference by enabling ultrafast token generation for highly responsive agentic systems.
Agentic systems can generate massive volumes of tokens across hundreds or thousands of inference steps, making faster token generation critical for agents to reason, act and complete complex tasks in real time.
Vera Rubin NVL72 systems provide the most versatile training and inference platform for every AI factory. NVIDIA Groq 3 LPX extends the inference performance of Vera Rubin NVL72 by dramatically increasing the rate of token generation, providing premium user experiences for context-heavy workloads so agents can act at extreme speeds.
NVIDIA Groq 3 LPX is pushing the frontier of AI inference. It delivered a record 3,400 output tokens per second in Artificial Analysis benchmarking running Gemma 4 31B, an open source agentic model, with a 100,000-token context critical for agentic systems — the fastest performance ever recorded for the model.
Groq 3 LPX enables agentic tasks such as coding in minutes versus hours, providing 4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform.
Login to comment