Tuesday, August 25, 2026
HomeAI AgentsNVIDIA Claims Vera Rubin NVL72 Delivers 30x More AI-Agent Work Per Watt...

NVIDIA Claims Vera Rubin NVL72 Delivers 30x More AI-Agent Work Per Watt Than Blackwell Ultra

NVIDIA says its next-generation Vera Rubin NVL72 system can do up to 30 times more useful AI-agent work per watt of power than its current Blackwell Ultra platform. That is a striking number, and it comes at a pointed moment: AI agents are already consuming far more computing power than ordinary chatbots, and electricity — not chips — is becoming the hard limit on how much AI infrastructure companies can actually build.

The claim, published by NVIDIA on August 24, 2026, is backed by a real benchmark methodology rather than an internal-only test. But it’s also an early-silicon, company-reported figure for hardware that customers can’t yet buy at scale. Here’s what NVIDIA actually measured, what it’s being compared against, and what still needs independent confirmation.

What NVIDIA Announced

NVIDIA’s claim centers on two numbers: Vera Rubin NVL72 delivers up to 30 times higher throughput per megawatt, and up to 35 times lower cost per million tokens, compared with NVIDIA’s current Blackwell Ultra system, known as GB300 NVL72. Both figures apply specifically to agentic AI workloads — the multi-step, tool-using tasks that have become the fastest-growing category of AI use.

To generate the numbers, NVIDIA used AgentX, an open-source benchmark built by the analysis firm SemiAnalysis as part of its InferenceX suite. Rather than testing a single prompt-and-response exchange, AgentX replays recorded real-world agentic coding sessions, preserving the growing context, tool calls, and sub-agent spawning that make agent workloads so much heavier than ordinary chat. That’s a meaningful detail: it means NVIDIA’s number is grounded in a benchmark methodology designed and published by an outside firm, not one NVIDIA built in-house — though NVIDIA still ran and reported the results itself.

It’s worth being precise about the comparison NVIDIA is making. The 30x figure measures Vera Rubin NVL72 against GB300 NVL72 — NVIDIA’s current-generation Blackwell Ultra rack — not against the older H100 or the earlier B100/B200 chips. That matters, because a separate figure in the same release states that GB300 NVL72 delivers up to 80 times higher throughput per megawatt than the still-older H200 NVL8 (Hopper generation) on large mixture-of-experts models. These are two different comparisons involving different hardware pairs, and it’s easy to see how they could get blurred together in casual reading.

Why Agentic AI Needs a New Efficiency Metric

The efficiency framing only makes sense once you understand how differently AI agents consume compute compared with a standard chatbot exchange.

According to data from OpenRouter, an AI model routing platform that has analyzed more than 100 trillion tokens of real-world usage, a single agentic request consumes roughly 15 times the tokens of an ordinary chat request. Separate OpenRouter reporting has also tracked agentic token volume growing roughly 14-fold since February 2026, even as human chat usage grew only a few times over the same period.

The reason is structural. A chatbot answers a question and stops. An AI agent doing real work — researching a company for an investment decision, for example — queries databases, searches filings, hands off sub-tasks to other agents, and keeps reasoning until the job is actually done. Each of those steps adds to a growing context that has to be carried forward into the next step, and long-context handling of that kind is expensive in a way that single-turn chat simply isn’t.

That is the underlying pressure NVIDIA is responding to. As agentic workloads scale, data centers increasingly run up against fixed power budgets rather than fixed chip supply — which is why the industry’s preferred metric is shifting from raw computing speed toward how much useful work a system can extract from each watt of electricity.

Two Different Multipliers — Don’t Mix Them Up

ComparisonMultiplierHardware pairWhat it measures
Vera Rubin NVL72 vs. GB300 NVL72Up to 30xNext-gen Rubin vs. current Blackwell UltraThroughput per megawatt on agentic coding workloads (AgentX benchmark)
GB300 NVL72 vs. H200 NVL8Up to 80xCurrent Blackwell Ultra vs. older Hopper generationThroughput per megawatt on large mixture-of-experts models

Both numbers come from NVIDIA’s same August 24 release, and both are legitimate efficiency claims — but they describe different generational jumps. The 30x figure is the one that applies to the still-unreleased Vera Rubin platform; the 80x figure describes a comparison between two generations of hardware that are both already shipping.

What’s Actually Inside Vera Rubin NVL72

Vera Rubin NVL72 isn’t a single chip — it’s a rack-scale system built from seven co-designed components: the Vera CPU, the Rubin GPU, an NVLink 6 switch, a ConnectX-9 network card, a BlueField-4 data processing unit, a Spectrum-6 Ethernet switch, and a newly integrated Groq 3 inference chip, acquired through NVIDIA’s purchase of Groq. A full NVL72 rack packs 72 Rubin GPUs and 36 Vera CPUs, cooled entirely by liquid rather than air.

The most consequential spec change is in memory. Rubin’s GPUs use HBM4 offering up to 288GB of capacity and up to 22 terabytes per second of bandwidth per GPU — roughly 2.75 times the bandwidth of Blackwell Ultra’s HBM3e. That’s a bigger jump than the compute increase, which NVIDIA and independent analysts put closer to 2.33 times, suggesting NVIDIA prioritized memory bandwidth over peak FLOPS this generation — a sensible choice, since agent workloads tend to be limited by how fast data moves rather than raw arithmetic.

The Power Trade-Off Behind “Work Per Watt”

It’s important to be clear about what “work per watt” does and doesn’t mean: Vera Rubin NVL72 does not use less electricity than Blackwell Ultra. It uses more.

Independent estimates converge on a Vera Rubin NVL72 rack drawing somewhere in the range of 180 to 230 kilowatts, compared with roughly 120 to 142 kilowatts for a GB300 or GB200 NVL72 rack — about 1.5 times the power draw of the prior generation. Per-GPU power consumption is reported to rise from around 1,400 watts on Blackwell Ultra to roughly 2,300 watts on Rubin. NVIDIA hasn’t published a single definitive rack-level power figure, so these numbers should be read as a converging range from multiple industry analyses rather than an official spec sheet.

The efficiency claim, in other words, isn’t about consuming less power — it’s about getting dramatically more useful output from each watt consumed, even as total consumption per rack climbs. That distinction matters for anyone actually planning data center capacity, since higher rack density means facilities also need upgraded liquid cooling and power delivery infrastructure to accommodate the new systems.

SpecVera Rubin NVL72GB300 NVL72 (Blackwell Ultra)AMD Helios (MI455X)
GPUs per rack727272
Memory typeHBM4HBM3eHBM4
HBM capacity (per rack)~20.7 TBNot consistently published in sources reviewed~31 TB (AMD claims ~50% more than Rubin)
Memory bandwidth (per GPU)Up to 22 TB/s~8 TB/sAMD claims ~6% more bandwidth than Rubin
Estimated rack power draw~180–230 kW~120–142 kWNot publicly disclosed in sources reviewed
Peak FP4 compute (rack)Not consistently published at rack level in sources reviewedNot consistently published at rack level in sources reviewedAMD claims ~15% more peak FP4 than Rubin
AvailabilityShipments guided for Q3 2026Currently shippingTargeting H2 2026

How Independent Is This Data?

This is the question that matters most for judging the 30x headline.

The clearest outside data point comes from CoreWeave, the cloud partner that completed the first Vera Rubin NVL72 rack bring-up on June 1, 2026. In July, CoreWeave published its own measured results using the DeepSeek-R1 reasoning model rather than the AgentX suite, and found Vera Rubin NVL72 delivering roughly 10 times more throughput per megawatt than GB300 NVL72 — a real but notably smaller gain, measured on a different workload.

SemiAnalysis, the firm behind the AgentX benchmark NVIDIA used, separately published its own throughput analysis of Rubin against Blackwell in July 2026. Against GB300 NVL72 specifically, it found Rubin’s advantage widening from roughly 2 times at moderate interactivity to as much as 5.4 times at the highest interactivity levels tested — a wide range, but still well short of 30x.

None of this makes NVIDIA’s 30x figure wrong — NVIDIA itself calls the numbers early results that should keep changing as software optimization continues. It does mean 30x sits at the high end of measured results, achieved under specific conditions, rather than a multiplier customers should expect across every workload.

When Can You Actually Get Vera Rubin NVL72?

Vera Rubin has moved through several stages of readiness over 2026, but it is not yet generally available to customers.

NVIDIA CEO Jensen Huang announced the platform was in full production at CES in January 2026, and detailed its full seven-chip architecture at GTC in March. CoreWeave’s rack bring-up in June marked the first time a complete system had been powered on and validated at production scale. NVIDIA has guided that broader partner shipments will begin in the third quarter of the 2026 calendar year, with named early partners including AWS, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, CoreWeave, Lambda, Nebius, and Nscale. Analysts at Morgan Stanley have estimated an initial shipment volume of roughly 150,000 units, representing about $9 billion in revenue, according to a pre-earnings analysis — an independent estimate, not a confirmed NVIDIA figure.

Some access already exists in limited form: Google Cloud has offered preview access to Rubin-based instances since the second quarter of 2026, but only in select regions and through an application process, not as a generally available product. “Full production” describes NVIDIA’s manufacturing status, in other words — it is a separate thing from broad commercial availability, which is still ahead.

DateMilestone
January 2026NVIDIA declares Vera Rubin platform in full production (CES)
March 2026Full seven-chip architecture detailed (GTC)
Q2 2026Google Cloud opens limited preview access in select regions
June 1, 2026CoreWeave completes first full rack bring-up and validation
July 2026CoreWeave and SemiAnalysis publish independent throughput measurements
August 24, 2026NVIDIA publishes the 30x/35x efficiency claims
August 26, 2026NVIDIA reports Q2 FY2027 earnings
Q3 2026 (guided)Broader partner shipments begin

Where Vera Rubin Sits Against the Competition

NVIDIA isn’t the only company chasing efficiency gains for agentic AI, and the competitive landscape has shifted since Vera Rubin was first announced.

AMD launched its own rack-scale competitor, Helios, built around 72 Instinct MI455X GPUs, at its Advancing AI event in mid-2026, with self-reported spec advantages summarized in the table above (source). Microsoft Azure and Anthropic have both announced plans to deploy Helios systems, with Anthropic committing to up to 2 gigawatts of AMD capacity — a reminder that major AI labs are pursuing multi-vendor strategies rather than relying on a single chipmaker.

Google’s TPU v7, known as Ironwood, has been generally available since late 2025 and is used heavily for Gemini workloads, while Amazon’s Trainium3 chips power AWS’s custom silicon push. Both use fundamentally different software ecosystems than NVIDIA’s CUDA platform, making direct efficiency comparisons difficult.

None of these competitors have been tested head-to-head against Vera Rubin NVL72 using the same benchmark. Every efficiency claim in this space — NVIDIA’s included — comes from a vendor choosing its own comparison conditions. No neutral third party has yet run an identical agentic workload across NVIDIA, AMD, Google, and AWS hardware, so cross-vendor superiority claims deserve some caution until that kind of independent test exists.

Business Stakes — Why This Lands Right Before Earnings

The timing is not incidental. NVIDIA reports second-quarter fiscal 2027 earnings on August 26, 2026 — one day after this data was published — and Wall Street has flagged Vera Rubin’s ramp as a key thing to watch. GF Securities analyst Jeff Pu recently raised his NVIDIA price target, citing rising orders for Vera Rubin from major cloud providers.

Some pre-earnings reports describe SpaceX committing to use Vera Rubin exclusively and Anthropic entering a $10 billion computing deal tied to NVIDIA infrastructure. Neither claim has been confirmed through a primary company filing reviewed for this article. Notably, Anthropic has also separately committed to deploying AMD’s Helios platform — a reminder that even major AI labs are hedging across chip vendors rather than committing to one supplier.

Limitations and Open Questions

Several open questions remain. NVIDIA itself describes the current results as preliminary, expecting performance to keep changing as software optimization continues on both platforms. The higher power density of Vera Rubin racks will also require upgraded liquid cooling and power delivery infrastructure, which could affect how quickly the efficiency gains translate into installed capacity.

One additional signal worth watching, not treating as settled: a chip-testing partner, King Yuan Electronics, reportedly lowered its guidance and deferred some AI-related revenue into later quarters — which some analysts attribute to longer qualification cycles rather than weakening demand. That distinction remains unresolved.

What to Watch Next

NVIDIA’s Q2 FY2027 earnings call on August 26, 2026, is the most immediate event to watch. Investors and industry analysts are expected to press for specifics on Q3 shipment volumes, customer uptake, and whether Blackwell Ultra demand holds steady while customers wait for Rubin. That call had not yet taken place at the time of this reporting, so its contents remain unknown.

Conclusion

NVIDIA’s 30x efficiency claim is grounded in a real, named benchmark and a genuine architectural shift toward memory bandwidth over raw compute — one that tracks with how agent workloads actually consume tokens. But it remains a company-reported result from early-silicon testing, measured against NVIDIA’s own prior generation under specific high-interactivity conditions rather than as a flat multiplier. The closest independent measurement, from CoreWeave, found a smaller but real gain of roughly 10 times on a different workload. Vera Rubin NVL72 is still in production ramp, not broad availability, with partner shipments expected around the third quarter of 2026. Whether the 30x figure holds at scale should become clearer as shipments begin and NVIDIA’s upcoming earnings call adds detail.

FAQ

What does NVIDIA’s “30x more work per watt” claim actually mean? It means NVIDIA measured Vera Rubin NVL72 completing up to 30 times more useful agentic AI work — using a benchmark that replays real coding-agent sessions — for the same amount of electricity as its current Blackwell Ultra (GB300 NVL72) system. It does not mean the new hardware uses less power overall; it uses more power per rack, but produces far more output per watt consumed.

Is Vera Rubin NVL72 compared to the H100, B200, or GB300? The 30x figure compares Vera Rubin NVL72 to GB300 NVL72, NVIDIA’s current Blackwell Ultra platform. A separate 80x figure in the same release compares GB300 NVL72 to the older H200 NVL8 (Hopper generation) — a different comparison entirely.

When will Vera Rubin NVL72 actually be available to customers? NVIDIA declared the platform in full production in January 2026, and a cloud partner, CoreWeave, completed the first full rack bring-up in June 2026. Broader shipments to partners including AWS, Google Cloud, Microsoft Azure, and others are guided to begin around the third quarter of 2026. As of this writing, it is not yet generally available to customers at scale.

Has anyone outside NVIDIA verified the 30x efficiency figure? Not at that exact magnitude. The closest independent data point is CoreWeave’s own July 2026 test, which measured roughly 10 times higher throughput per megawatt using a different benchmark model. NVIDIA describes its 30x figure as an early result that may shift as software optimization continues.

Why do AI agents use so many more tokens than chatbots? According to OpenRouter data, a single agentic request consumes roughly 15 times the tokens of an ordinary chat request, because agents perform multi-step tasks — querying data sources, calling tools, spawning sub-agents — and carry a growing context forward through each step, rather than answering once and stopping.

How does Vera Rubin NVL72 compare to AMD’s Helios rack? AMD’s own comparisons claim its Helios platform, built around 72 Instinct MI455X GPUs, offers somewhat higher peak compute, more memory capacity, and better cost-per-token than Vera Rubin NVL72. These figures come from AMD’s own testing, however, and no neutral third party has benchmarked the two platforms against each other using the same workload.

Does higher efficiency mean Vera Rubin NVL72 uses less power? No. Estimates suggest a Vera Rubin NVL72 rack draws roughly 180 to 230 kilowatts, compared to about 120 to 142 kilowatts for a Blackwell Ultra rack — about 1.5 times more power. The efficiency gain comes from doing far more useful work within that larger power budget, not from reducing consumption.

RELATED ARTICLES
- Advertisment -

Most Popular