Wednesday, August 26, 2026
HomeAI InfrastructureNvidia Isn't Just Selling GPUs Anymore — It's Selling the Glue That...

Nvidia Isn’t Just Selling GPUs Anymore — It’s Selling the Glue That Holds AI Factories Together

For years, the story of Nvidia’s dominance was simple: it made the best AI chips, and everyone who wanted to train or run large models had to buy them. That story is getting more complicated. Amazon, Google, Microsoft, and Meta are all now designing their own AI accelerators, hoping to cut their dependence on Nvidia hardware. Nvidia’s response, laid out again in a company blog post this week, is not to compete chip-for-chip with each of them. It’s to make sure that whatever chips they build still have to run through Nvidia’s plumbing.

That plumbing is called NVLink Fusion, and it’s the centerpiece of what Nvidia calls the “AI factory” — a framing the company has used since 2024 to describe data centers that convert electricity and data into AI output around the clock. The pitch is straightforward: hyperscalers can keep designing their own custom chips, but if they want those chips to perform at scale, they’ll still need to plug into Nvidia’s networking, software, and rack infrastructure.

NVLink Fusion itself isn’t new. Nvidia introduced it in May 2025, and this week’s post is best understood as a restatement and expansion of that strategy — not a new product launch. What has changed is the framing: Nvidia is increasingly asking the industry to judge AI infrastructure not by how fast an individual chip is, but by system-level economics like tokens produced per watt of power and cost per token generated.

NVLink Fusion is Nvidia’s technology for letting outside companies’ custom accelerators — generally referred to across the industry as “XPUs” — connect directly into Nvidia’s high-speed NVLink network, rather than being isolated on their own separate infrastructure.

The current generation, NVLink 6, can link up to 72 accelerators together in a single high-bandwidth domain, with each accelerator getting up to 3.6 terabytes per second of connectivity, according to Nvidia. The company says future versions of the roadmap will support domains of up to 1,152 accelerators, along with co-packaged optics for even higher-speed connections. A related piece of the technology, NVLink-C2C, handles chip-to-chip links, allowing a custom CPU or accelerator to connect directly to Nvidia GPUs with a shared memory pool.

DateDevelopment
May 2025NVLink Fusion unveiled at Computex; initial partners include Marvell, MediaTek, Alchip, Astera Labs, Synopsys, Cadence
November 2025Arm adds NVLink Fusion support to its Neoverse platform
December 2025AWS confirms NVLink Fusion adoption for Trainium4
February 2026Meta reportedly shuts down in-house Olympus chip project
March 2026Nvidia invests $2 billion in Marvell alongside expanded NVLink Fusion partnership
April 2026UALink Consortium releases version 2.0 specification
August 2026Nvidia reaffirms the strategy in its “AI factory” blog post

The physical connection is only part of it. Nvidia bundles NVLink Fusion with its MGX rack architecture — essentially a standardized blueprint for how servers, cooling, and power are laid out — plus a software stack that includes its NCCL communications library, the Dynamo and NIXL tools for splitting up AI workloads, and Mission Control, a platform for managing and monitoring large clusters. The idea is that a company adopting NVLink Fusion doesn’t just get a faster cable — it gets access to Nvidia’s entire operational playbook for running AI infrastructure at scale, which Nvidia says reduces the risk and cost of building a rack-scale system from scratch.

It’s worth pausing on the term “XPU” itself, since it shows up constantly in this conversation. It isn’t an Nvidia invention or a specific chip category. Across the industry, “XPU” functions as a generic placeholder — a way of referring to any specialized processor, whether it’s a GPU, a custom AI accelerator, or something else, without committing to a particular architecture. Nvidia uses it the same way everyone else does.

Who Has Actually Signed On

This is where the picture gets more nuanced than the “everyone is joining Nvidia’s ecosystem” narrative might suggest.

The clearest, most concrete example is Amazon Web Services. AWS is adopting NVLink Fusion to help build out infrastructure around its new Trainium4 chip, combining it with Nvidia’s MGX rack design, AWS’s own Graviton CPUs, and its Elastic Fabric Adapter networking. This is a real, named production commitment — not a hypothetical.

Beyond AWS, Nvidia has built a wide roster of silicon and IP partners: Marvell, MediaTek, Alchip Technologies, Astera Labs, Synopsys, Cadence, Fujitsu, Qualcomm, Intel, and Samsung Foundry have all signed on to design or manufacture chips compatible with NVLink Fusion. In November 2025, Arm added NVLink Fusion support to its Neoverse platform — the design basis for a huge share of custom cloud chips, including Google’s Axion processors, Amazon’s Graviton4, and Microsoft’s Azure Cobalt 100 CPUs.

That last point is often where coverage overstates things. Arm’s Neoverse supporting NVLink Fusion means that companies building on Neoverse could connect their designs to Nvidia’s fabric. It is not the same as Google or Microsoft announcing they will actually do so for their AI accelerators. As of this writing, there is no public confirmation that Google’s TPUs or Microsoft’s Maia chips use NVLink Fusion. The compatibility exists at the architectural level; the hyperscaler commitment, for those two companies, does not — at least not yet, and not publicly.

CompanyRelationship to NVLink Fusion
AWSConfirmed production adopter — using it for Trainium4/MGX rack deployment
Marvell, MediaTek, Arm, Fujitsu, Qualcomm, Intel, Samsung FoundryConfirmed silicon/IP design partners (ecosystem-level, not end-customer chip adoption)
GoogleNo public confirmation of NVLink Fusion adoption for TPUs; Arm Neoverse (used in Google’s Axion CPUs) supports NVLink Fusion at the IP level
MicrosoftNo public confirmation of NVLink Fusion adoption for Maia chips; Azure Cobalt 100 CPU built on Neoverse, which supports NVLink Fusion at the IP level
MetaNo public confirmation of NVLink Fusion adoption; in-house Olympus chip project reportedly discontinued in February 2026

Nvidia has also worked to keep the partner ecosystem close. In March 2026, it invested $2 billion in Marvell alongside an expanded NVLink Fusion partnership, a move that industry analysts have interpreted as an effort to keep one of the two dominant custom-chip design houses — Marvell and Broadcom design most of the industry’s custom AI silicon — tied to Nvidia’s ecosystem rather than drifting entirely toward open, non-Nvidia standards.

The Pitch: Tokens per Watt, Not Just Raw Speed

Nvidia’s broader argument is that comparing AI chips on raw speed or price per unit misses the point. What matters, the company says, is how much useful AI output — measured in tokens, the small chunks of text or data that language models generate — a system can produce per watt of electricity and per dollar spent.

Nvidia backs this up with striking numbers: it says its Blackwell architecture cuts the cost of generating a million tokens by 15 times compared with the prior generation. For its most advanced current variant, Blackwell Ultra — sold as the GB300 NVL72 system — Nvidia claims up to 50 times more tokens per megawatt of power and up to 35 times lower cost per token than its 2022-era Hopper chips.

Independent testing offers a more mixed picture. Benchmarks published by the research firm SemiAnalysis, using its InferenceMAX testing framework, found real efficiency gains generation-over-generation — for example, roughly a threefold improvement in tokens processed per megawatt moving from Nvidia’s H100 to its B200 chip on one workload. But those same benchmarks also found much narrower differences in some head-to-head comparisons between vendors: in one test, Nvidia’s Blackwell chips came out only about 20% more power-efficient than AMD’s competing CDNA4 architecture — a meaningful edge, but a far cry from the order-of-magnitude gaps sometimes implied in marketing materials.

There’s also a fair question about whether tokens per watt is a neutral way to measure chip performance at all. Because the metric depends on the entire system — the chip, the networking, the software stack, and how well they’re tuned together — it tends to favor a company like Nvidia that controls all of those pieces itself. A competitor selling a bare chip without Nvidia’s surrounding software ecosystem may look worse on this measure even if the chip itself is competitive, simply because it isn’t running on Nvidia’s optimized stack. That doesn’t make the metric meaningless, but it does mean readers should treat Nvidia’s specific multiples as company-reported figures rather than settled, vendor-neutral facts.

The Competitive Response Is Real

Nvidia isn’t operating in a vacuum. A separate, genuinely open effort to build alternatives to NVLink has been underway for two years and is starting to produce actual hardware.

The UALink Consortium — backed by AMD, Broadcom, Google, Microsoft, Meta, Intel, and Cisco, among others — has developed an open standard for connecting AI accelerators that doesn’t require Nvidia’s proprietary technology. The group released version 2.0 of its specification in April 2026, and AMD has said its MI400 chips, built for the “Helios” rack system it previewed at CES in January 2026, will support the standard more broadly in the second half of 2026.

Separately, Broadcom has been building its own answer using standard Ethernet rather than a specialized protocol. Its Tomahawk Ultra switch chip, which began shipping in mid-2025, is built specifically for connecting large numbers of AI accelerators with very low latency, and Broadcom says it can support rack configurations of up to 1,024 accelerators — well beyond the 72-accelerator domain of Nvidia’s current NVLink Switch. A companion effort, the Ultra Ethernet Consortium, is developing complementary standards for connecting AI hardware across entire data centers, and demonstrated interoperability of its specification with Broadcom hardware at a networking conference in early 2026.

The competitive landscape here is genuinely tangled rather than a clean two-sided fight. Broadcom, for instance, is simultaneously one of Nvidia’s biggest potential threats through its Ethernet-based alternatives and, according to industry analysts, a contributor to parts of the NVLink Fusion arrangement itself — supplying components that make Nvidia’s own open-ecosystem push possible. It’s less “Nvidia versus Broadcom” and more a set of overlapping alliances where companies compete and cooperate with Nvidia at the same time.

As of mid-2026, though, most of these open alternatives are still earlier in their commercial rollout than NVLink Fusion, which has a year-plus head start and a working, revenue-generating deployment in AWS’s Trainium4 buildout.

InterconnectBackersDomain sizeStatus (as of mid-2026)
NVLink Fusion (Nvidia)Nvidia + partner ecosystem (Marvell, MediaTek, Arm, Fujitsu, Intel, Samsung Foundry, others)72 accelerators (NVLink 6); roadmap to 1,152Shipping; one confirmed hyperscaler deployment (AWS Trainium4)
UALinkAMD, Broadcom, Google, Microsoft, Meta, Intel, CiscoUp to 1,024 accelerators (spec)Version 2.0 released April 2026; AMD MI400 hardware expected more broadly in H2 2026
Broadcom Tomahawk Ultra / SUEBroadcomUp to 1,024 acceleratorsShipping since mid-2025
Ultra Ethernet Consortium (UEC)Broadcom, Keysight, and other membersData-center-wide (scale-out, not scale-up)Interoperability demonstrated in 2026; complements UALink rather than replacing it

Note: Figures are company-reported specifications; independent head-to-head benchmarks between these systems are not yet available.

Why Nvidia Is Doing This

The strategic logic isn’t complicated. Custom AI chips from Google, Amazon, Microsoft, and Meta are growing fast — one industry estimate puts the growth rate at roughly 45% a year — and some analysts project Nvidia’s share of the AI inference chip market could fall from more than 90% today to somewhere between 20% and 30% by 2028. That’s a projection, not a certainty, but it captures the pressure Nvidia is under.

NVLink Fusion is Nvidia’s way of making sure that even in a data center full of a rival’s custom chips, Nvidia still has a role — supplying the switch, the networking card, or the CPU that ties the whole system together. As one industry analysis put it, the goal is for Nvidia to participate in every hyperscaler’s infrastructure, whether its own chips make up 10% or 90% of the compute in the room.

It’s also worth noting that not every hyperscaler’s custom-silicon ambitions are going smoothly, which complicates any simple story about inevitable independence from Nvidia. Meta reportedly shut down its most advanced in-house training chip project, code-named Olympus, in February 2026 after concluding its software wasn’t stable enough and its design too complex to manufacture at scale — a reminder that building competitive AI silicon is difficult even for companies with enormous resources.

What’s Still Unknown

Several important details remain unresolved. Nvidia has not publicly disclosed pricing or licensing terms for NVLink Fusion, so it’s unclear how the economics compare to a hyperscaler building its own interconnect from scratch. There are also no independent, head-to-head benchmarks yet comparing a full NVLink Fusion rack against an equivalent UALink or Broadcom-based system, largely because competing hardware is still limited in availability. And there’s no public confirmation, one way or the other, of whether Google or Microsoft plan to use NVLink Fusion for their own custom chips going forward.

What It Means Now

The clearest takeaway from Nvidia’s current position is that the company’s competitive strategy has shifted in an important way: it no longer needs to win every chip design competition to remain central to AI infrastructure. By opening its interconnect to outside chips, Nvidia is betting that the network and system layer — not the chip itself — is where lasting advantage lives.

That bet has real traction, evidenced by a working AWS deployment and a broad partner ecosystem. But it is not yet the industry-wide consolidation some coverage suggests. Only one hyperscaler has confirmed production use, several credible open alternatives are advancing in parallel, and the efficiency numbers Nvidia promotes most heavily come with important caveats when checked against independent testing.

What to watch next: whether Google or Microsoft make any public statement about NVLink Fusion for their own chips, how quickly AMD’s UALink-compatible MI400 hardware actually ships in the second half of 2026, and whether independent benchmarks emerge comparing Nvidia’s ecosystem directly against the open alternatives now reaching the market.


FAQ

What is NVLink Fusion, in plain terms? It’s Nvidia’s technology that lets other companies’ custom AI chips connect to Nvidia’s high-speed networking and rack infrastructure, instead of requiring an entirely separate, from-scratch system.

What does “XPU” mean, and is it an Nvidia term? No — it’s a generic industry term used across the semiconductor world as a placeholder for “some kind of processing unit,” whether that’s a GPU, a custom AI accelerator, or another specialized chip. Nvidia uses it the same way its competitors do.

Which companies have actually adopted NVLink Fusion? AWS is the only hyperscaler with a confirmed production commitment, using it for its Trainium4 chip. A broad set of chip designers and manufacturers — including Marvell, MediaTek, Arm, Fujitsu, Qualcomm, Intel, and Samsung Foundry — have signed on as ecosystem partners, but that’s different from a hyperscaler committing its own AI chip to the platform. There’s no public confirmation that Google or Microsoft plan to use it for their TPU or Maia chips.

What is an “AI factory,” and how is it different from a regular data center? It’s Nvidia’s term, in use since 2024, for data centers designed to continuously generate AI output — measured in tokens — rather than simply store and retrieve information the way traditional data centers do.

What does “tokens per watt” mean, and why is Nvidia emphasizing it? It measures how much AI output a system produces for each watt of electricity it consumes. Nvidia is emphasizing it because power availability is becoming a major constraint on AI infrastructure, and the company argues this system-level metric is more meaningful than chip-level specs alone. Independent testing shows real gains generation over generation, but often smaller than Nvidia’s headline multiples suggest, particularly in some cross-vendor comparisons.

What are the open alternatives to NVLink Fusion? The UALink Consortium (backed by AMD, Broadcom, Google, Microsoft, and Meta) and Broadcom’s own Ethernet-based products, including its Tomahawk Ultra switch, both aim to offer non-Nvidia ways to connect large numbers of AI chips together. A related industry group, the Ultra Ethernet Consortium, is developing complementary standards for connecting AI hardware across entire data centers.

Does NVLink Fusion mean Nvidia GPUs are becoming less important? Not necessarily. It means Nvidia is hedging against a future where its own chips make up a smaller share of AI infrastructure, by making sure its networking and software remain valuable even in systems built around competitors’ chips.

RELATED ARTICLES
- Advertisment -

Most Popular