Nvidia's Groq 3 LPX Targets the Agent Inference Bottleneck at Rack Scale
Nvidia enters full production on a dedicated inference accelerator that disaggregates token generation from context processing, directly threatening standalone inference chip rivals.
2. Nvidia's Groq 3 LPX Targets the Agent Inference Bottleneck at Rack Scale
Announced at Hot Chips 2026 on August 24, Nvidia's Groq 3 LPX inference accelerator has entered full production as a purpose-built extension to the Vera Rubin NVL72 rack-scale platform. The chip is designed to offload token-generation workloads from Vera Rubin's GPUs, allowing a single rack deployment to connect up to 256 LP30 accelerators over Nvidia's high-bandwidth interconnects. Nebius Group N.V. signed on as the first committed customer. In benchmark tests by Artificial Analysis, Groq 3 LPX outputs 3,400 tokens per second running Gemma 4 31B with a 100,000-token context window, a result Nvidia claims is four times faster than rival platforms on latency-sensitive workloads.
The strategic move here is disaggregation. Agentic workloads generate massive decode latency because today's GPU clusters handle context ingestion and token generation on the same silicon. Groq 3 LPX splits those two functions: Vera Rubin GPUs absorb context, LPX accelerators handle decode. That eliminates the throughput-versus-latency tradeoff that has made scaling agent pipelines painful. The underlying technology comes from Groq Inc., the inference-focused chip startup Nvidia licensed for $20 billion in December 2025, also bringing aboard founder Jonathan Ross and president Sunny Madra. That acquisition price signals how seriously Nvidia views inference-only silicon as the next margin battleground, and it puts direct pressure on standalone inference providers like Cerebras and on cloud inference services built around commodity GPU clusters.
The broader pattern is a consolidation of the AI compute stack under one vendor. Nvidia now controls training silicon, the interconnect fabric, and dedicated inference acceleration in a single rack-scale configuration. Customers who standardize on Vera Rubin NVL72 plus Groq 3 LPX will find it increasingly difficult to swap in third-party inference chips without sacrificing the tight GPU-to-LPU bandwidth that makes the performance numbers real. Watch whether hyperscalers accept that lock-in or accelerate internal inference silicon programs as a direct response.
Source: Nvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents