Nvidia said this week that Groq 3 LPX, its dedicated AI inference accelerator, has entered full production — the first hard milestone since the chipmaker's roughly $20 billion Groq deal.
According to CNBC, Nvidia says racks built around the Groq chips will be online this year, a timeline the network frames as evidence of how important low-latency inference has become to the AI business. Techmeme, citing Mike Wheatley at SiliconANGLE, reports that cloud provider Nebius has signed on as the first Groq 3 LPX customer, and that SpaceXAI will adopt Nvidia's Vera CPUs. Nvidia's own newsroom describes the SpaceXAI deployment as covering both Vera CPUs and the Vera Rubin platform, with Quiver Quantitative and Stock Titan noting the company's stated ambition to extend that computing into orbit through a project called Starmind.
On its blog, Nvidia positions the launch as an extension of Vera Rubin NVL72 rather than a standalone product, arguing that "the next era of AI inference won't be defined by a single breakthrough chip, network or system" but by how the layers of an AI factory work together. The company claims Vera Rubin NVL72 delivers up to 30x more work per watt, and cites OpenRouter data showing that agentic AI workloads consume 15 times more tokens than a simple chat request. Wccftech's write-up goes further, billing the combination as producing the fastest token generation speeds ever recorded.
Those speed and efficiency figures come from Nvidia and outlets summarizing Nvidia, not independent benchmarks, so they should be read as vendor claims for now.
Why it matters: the money in AI is shifting from training models to running them constantly for millions of users, and AI agents that call tools and sub-agents burn far more tokens than chatbots do — so whoever makes inference fast and cheap per watt controls the economics of the next phase.