Tuesday, August 25, 2026
HomeAI InfrastructureGroq 3 LPX Enters Full Production as Nvidia Extends Vera Rubin for...

Groq 3 LPX Enters Full Production as Nvidia Extends Vera Rubin for AI Agents

Nvidia has moved a key piece of its AI inference strategy from blueprint to shipping hardware. The company announced at the Hot Chips 2026 conference on August 24 that Groq 3 LPX, a chip built from technology it licensed from AI accelerator startup Groq, has entered full production and is now being integrated into Nvidia’s Vera Rubin NVL72 rack-scale platform, alongside new networking and interconnect technology aimed at one problem: making AI agents respond faster.

The announcement marks the first concrete commercial step in a deal that has reshaped the AI chip landscape since it was struck in December 2025. Cloud provider Nebius Group has signed on as the first customer.

What Nvidia Announced

Nvidia says Groq 3 LPX — a rack of processors purpose-built for fast token generation — is now in full production and is being woven into the Vera Rubin NVL72 system using Nvidia’s Spectrum-X networking and a newly introduced interconnect standard called NVLink Fusion.

This is not a new corporate deal — it’s a production milestone built on an agreement Nvidia and Groq announced in December 2025. Nvidia first showed how Groq’s technology would fit into Vera Rubin at its GTC conference in March 2026; the August announcement confirms the resulting chip is now shipping to at least one named customer, Nebius, through its Nebius Token Factory inference service.

Why It Matters: The Agentic AI Latency Problem

Running an AI “agent” is a different job than simply answering a question. Agents plan, call tools, check their own work and loop through many small reasoning steps before finishing a task — and if any of those steps is slow, the whole agent feels sluggish.

Nvidia’s GPUs are built to move enormous amounts of data through a model at once, a strength for training and processing long inputs. Groq’s chip architecture, the LPU, or Language Processing Unit, was designed for something narrower: generating tokens one after another as quickly and predictably as possible — exactly what agentic workloads reward. By combining the two, Nvidia is arguing that Vera Rubin can handle both halves of the job, ingesting and reasoning over large context, then generating responses with minimal delay, inside a single rack-scale system.

How Nvidia Got Groq’s Technology

The relationship between Nvidia and Groq is easy to misread, so it’s worth being precise about it.

[Timeline graphic: Dec 2025 Nvidia-Groq deal → Feb 2026 Warren/Wyden/Blumenthal FTC/DOJ referral request → March 2026 GTC unveiling of Vera Rubin → March 2026 Warren/Blumenthal letter to Nvidia → Aug 2026 Groq 3 LPX full production announcement.]

In December 2025, Nvidia agreed to pay roughly $20 billion in cash — reportedly its largest transaction ever — for a non-exclusive license to Groq’s inference chip patents and technology. As part of the agreement, several of Groq’s senior technical staff joined Nvidia, reportedly including the company’s CEO and president.

Groq itself was not sold, and it did not disappear. Its cloud service, GroqCloud, has continued to operate independently, and in May 2026 Groq was reported to be raising an additional $650 million from existing investors to grow that business. Analysts and journalists have widely described the Nvidia arrangement as a “reverse acquihire” — a deal structured to give one company most of the practical benefits of an acquisition without the formal process of buying the company outright.

Inside the Vera Rubin Platform

Nvidia describes Vera Rubin as seven chips across five rack-scale configurations: the Vera Rubin NVL72 compute rack, a Vera CPU rack, the new Groq 3 LPX inference rack, a Spectrum-6 networking rack, and a BlueField-4 storage and security rack.

The Vera Rubin NVL72 rack itself pairs 72 Rubin GPUs with 36 Vera CPUs, connected through Nvidia’s internal NVLink interconnect so the whole rack behaves like one very large processor. The Groq 3 LPX rack is described as containing 256 LPU processors, with 128GB of on-chip SRAM and 640 terabytes per second of internal bandwidth — figures aimed at showing why the chip can move data between its own processors fast enough to keep token generation flowing without delay.

Alongside the Groq integration, Nvidia introduced NVLink Fusion, a version of its interconnect technology that lets custom chips built by other companies connect directly into Nvidia’s rack-scale networking — meaning a cloud provider could build a rack mixing Nvidia GPUs with non-Nvidia processors and still link them via Nvidia’s high-speed fabric.

Nvidia and independent coverage agree on that broad purpose, but detailed technical specifications for NVLink Fusion itself — bandwidth, latency, or exact protocol details separate from Nvidia’s existing NVLink standard — have not been published. It’s better understood for now as a stated capability than a fully documented specification.

The Performance Claims: What’s Verified and What Isn’t

Nvidia’s central performance claim is that, in a benchmark run by the third-party testing service Artificial Analysis, a Vera Rubin system running the open-source Gemma 4 31B model produced 3,400 output tokens per second on a task involving a 100,000-token context — the kind of long, complex prompt agentic systems typically handle. Nvidia says that result is four times faster than the nearest alternative platform it compared it to, and separately claims up to 30 times higher token throughput per megawatt of power compared with its own previous-generation GB300 NVL72 system.

These are real numbers from a legitimate benchmarking service, but the specific test run was selected and published by Nvidia, not conducted independently by a neutral third party comparing all vendors under identical conditions. No outlet has yet published its own side-by-side replication, and there is no public comparison showing how Groq 3 LPX performs against earlier generations of Groq’s own chips. None of that means the numbers are wrong — it means readers should treat them as Nvidia’s own reported results until independent testing catches up.

Early Deployment and Customers

Nebius Group is the first company Nvidia has named as committed to deploying Groq 3 LPX, through its Nebius Token Factory inference offering. Separately, cloud provider CoreWeave has already put Spectrum-X Multiplane — the networking technology that connects Vera Rubin racks at scale — into production use.

Nvidia has said Groq-equipped Vera Rubin racks will be “online this year,” which points to a 2026 rollout rather than an immediate, broad general release. Pricing and a full list of customers have not been disclosed.

Competitive Landscape

Nvidia is not the only company chasing faster agentic inference. AMD has introduced its Instinct MI400 series of accelerators, built on its CDNA 5 architecture with HBM4 memory, positioning them as a broad rival to Vera Rubin on raw memory capacity and bandwidth.

AMD has also partnered with Cerebras, pairing its Helios racks with Cerebras’ wafer-scale processors for low-latency inference — a move industry reporting has explicitly described as AMD’s answer to Nvidia’s deal with Groq. AMD’s chips also support UALink, an open interconnect standard backed by AMD, Intel, Google, Meta and Microsoft, positioned as an alternative to Nvidia’s proprietary NVLink and NVLink Fusion approach.

Other inference-focused chipmakers, including SambaNova and Intel with its Gaudi line, compete in the same broad market, though neither has announced a specific response tied to this particular Nvidia-Groq milestone.

Nvidia Vera Rubin + Groq 3 LPXAMD Instinct MI400 + Cerebras
Core approachGPU rack (Vera Rubin NVL72) paired with a dedicated latency-optimized inference chip (Groq 3 LPX)GPU accelerator (MI400 series, CDNA 5, HBM4) paired with Cerebras wafer-scale engines for disaggregated inference
InterconnectProprietary NVLink 6 / NVLink-C2C, with NVLink Fusion allowing some third-party chips to connectUALink, an open interconnect standard backed by multiple companies
Stated targetAgentic AI workloads requiring fast, iterative token generationBroad AI training/inference, with the Cerebras partnership targeting low-latency disaggregated inference
Production statusGroq 3 LPX in full production as of August 2026MI400 series unveiled in 2026; Cerebras partnership announced mid-2026

Regulatory Scrutiny of the Nvidia-Groq Deal

The unusual structure of the Nvidia-Groq arrangement has drawn attention in Washington. In March 2026, Senators Elizabeth Warren and Richard Blumenthal sent a letter to Nvidia CEO Jensen Huang questioning whether the deal was designed to avoid the antitrust review that a formal acquisition would typically require, and warning that it could further entrench Nvidia’s dominance in AI chips.

That letter followed a broader February 2026 push in which Warren, Ron Wyden and Blumenthal called on the FTC and Department of Justice to examine a pattern of similar “reverse acqui-hire” deals across the industry, not limited to Nvidia.

This is a Congressional inquiry raising concerns and requesting information — not a formal government investigation, lawsuit, or any finding that the deal violated the law. As of the most recent information available, no formal enforcement action against the Nvidia-Groq deal has been confirmed; readers should treat this as an open, ongoing question rather than a settled legal matter.

Limitations and Open Questions

Several details about this announcement remain unclear. The full terms of Nvidia’s license with Groq — including its exact scope and duration — have not been made public. No independent organization has yet replicated Nvidia’s performance figures for Groq 3 LPX under controlled, vendor-neutral conditions. Nvidia’s own materials also reference an additional partner adopting its Vera CPUs for agentic workloads, tied to plans for future orbital deployment — a claim that current reporting cannot clearly tie to a single identified company, and should be read with caution.

Conclusion

What’s confirmed is straightforward: Groq 3 LPX, hardware Nvidia obtained through a licensing and talent deal with Groq in December 2025, is now in full production and shipping as part of the Vera Rubin platform, with at least one named customer. What remains open matters just as much — Nvidia’s headline performance numbers are its own, not yet independently confirmed, and the deal that made this hardware possible is still being questioned by lawmakers over whether it sidestepped standard antitrust review. Whether the performance claims hold up under independent scrutiny, and whether regulators take any formal action, are the two threads most worth watching next.

Frequently Asked Questions

Did Nvidia acquire Groq? No. Nvidia paid roughly $20 billion for a non-exclusive license to Groq’s chip technology and hired several of its senior staff, but Groq was not purchased outright. It continues to operate its own cloud business, GroqCloud, independently.

What is Groq 3 LPX and how is it different from a GPU? Groq 3 LPX is an inference-focused chip built from licensed Groq technology, designed to generate tokens quickly and predictably. Unlike GPUs, which are optimized for processing large amounts of data in parallel, it’s tailored specifically for the fast, sequential token generation that AI agents rely on.

What is NVLink Fusion? It’s a version of Nvidia’s interconnect technology that allows chips designed by other companies to connect into Nvidia’s rack-scale networking, giving hyperscalers a way to mix Nvidia and non-Nvidia hardware in the same system.

How fast is Groq 3 LPX, according to Nvidia’s benchmarks? Nvidia cites a benchmark showing 3,400 output tokens per second on a 100,000-token task using the Gemma 4 31B model, via testing service Artificial Analysis. This figure comes from Nvidia’s own selected test and has not yet been independently replicated.

Who is using Groq 3 LPX so far? Nebius Group is the first publicly named customer, planning to use it through its Nebius Token Factory service. CoreWeave has separately deployed related Spectrum-X networking technology in production.

Is the Nvidia-Groq deal under investigation? It has drawn a formal information request from Senators Elizabeth Warren and Richard Blumenthal over whether it was structured to avoid antitrust review. No formal government investigation or enforcement action against the deal has been confirmed.

When will Vera Rubin racks with Groq 3 LPX be available? Nvidia has said the racks will be operational “this year,” meaning sometime in 2026, but has not disclosed pricing or a full commercial availability timeline.

RELATED ARTICLES
- Advertisment -

Most Popular