Groq1 article

Groq

Articles

  • NVIDIA Groq 3 LPX Enters Full Production Delivering 3,400 Tokens per Second for Agentic Inference

    At Hot Chips 2026, NVIDIA announced that the Groq 3 LPX dedicated inference accelerator has entered full commercial production. Built as a workload-optimized extension to the Vera Rubin data center architecture, the accelerator addresses the decode bottlenecks inherent in multi-step agentic AI workloads by offloading token generation from primary GPU clusters onto specialized Language Processing Units (LPUs). In independent benchmarking conducted by Artificial Analysis, the Groq 3 LPX delivered

    1 min