Friday, September 18, 2026
Advertisement
HomeEnterprise AIEquinix Inference Exchange: Announced With NVIDIA and Together AI, Arriving Q1 2027

Equinix Inference Exchange: Announced With NVIDIA and Together AI, Arriving Q1 2027

0
(0)

Equinix announced Equinix Inference Exchange on 2 September 2026, a distributed AI inference offering built with NVIDIA and Together AI — and scheduled to become available in the first quarter of 2027, not today.

That gap between announcement and availability is the most important thing to understand about the news. Equinix has described an architecture, named its partners, and put a quarter on the calendar. What it has not described is where the service will run, what it will cost, which hardware it will run on, or how fast it will be. For enterprises trying to decide where their AI inference workloads should live, that combination makes Inference Exchange something to track rather than something to plan around.

What Equinix announced

The announcement came at Equinix Horizon, the company’s first customer and partner event in its 28-year history, held in Redwood City. Equinix framed Inference Exchange as an expansion of its existing collaboration with NVIDIA, joined by a new collaboration with Together AI. It was announced alongside Equinix Fabric One, a separate managed connectivity service.

Advertisement

The underlying argument is straightforward. As enterprise AI moves from experimentation into production, inference increasingly needs to run close to the users, data and applications it serves — across multiple clouds, providers and jurisdictions. Equinix’s pitch is that its interconnection footprint is already sitting in those places.

Equinix chief executive and president Adaire Fox-Martin described the resulting architectures as “neutral by design, open by default” — positioning language rather than a verified technical property, but a clear signal of how the company wants the offering understood against hyperscaler alternatives.

How the three layers fit together

Inference Exchange is a stack assembled from three companies, each contributing a distinct layer. All of it is described in future tense in Equinix’s materials.

LayerContributorWhat it provides
InfrastructureEquinixData centres, power, advanced cooling, day-two operations, and connectivity through Equinix Fabric to clouds, networks and AI providers
Compute designNVIDIAValidated Enterprise Reference Architectures and AI infrastructure underpinning the build
Inference platformTogether AIThe inference platform itself, with access to more than 200 open-source models, in both multitenant and dedicated single-tenant deployments

The multitenant and single-tenant distinction matters for buyers. Multitenant deployments are described as the shared-efficiency option; dedicated single-tenant environments are aimed at workloads that require reserved capacity or greater isolation. Both are planned, though neither has been priced or detailed.

One clarification worth making: the “more than 200 open-source models” figure describes Together AI’s platform as a whole. Equinix has not published a model list for Inference Exchange specifically, and it has not said whether the full catalogue will be reachable through the offering.

What an NVIDIA Enterprise Reference Architecture actually is

Because the NVIDIA layer is described only by programme name, it is worth explaining what that programme covers — while being clear this is general NVIDIA documentation, not an Inference Exchange specification.

NVIDIA’s Enterprise Reference Architectures are validated blueprints for building AI clusters. They begin with tested hardware configurations covering CPU, GPU and networking node patterns, along with cabling diagrams and infrastructure detail, and extend into a software reference stack currently centred on open-source Kubernetes together with NVIDIA AI Enterprise and NVIDIA Run:ai. They exist as a family of designs rather than a single blueprint.

Which member of that family Equinix intends to deploy — and therefore which generation of accelerators enterprises would actually be renting — has not been disclosed.

The three scenarios Equinix is targeting

Equinix named three enterprise scenarios for the offering:

Metro edge inference. Running inference physically closer to the users and data it serves, to support lower-latency AI experiences.

Open model migration. Providing a route for enterprises moving workloads off closed, proprietary models toward open-source alternatives, to control cost and reduce lock-in — reachable over the same interconnection fabric those enterprises already use for other providers.

Sovereign AI. Enabling workloads to run in locations that support data residency and sovereignty requirements for regulated industries.

The sovereign AI case carries an obvious tension. Data residency is a question about specific jurisdictions, and Equinix has not named a single launch location. Independent commentary has also noted that sovereignty in practice extends well beyond physical placement — encompassing identity, encryption key control, logging, and the handling of model weights, prompts and generated output. None of those controls have been specified for this offering.

Availability: preview now, Q1 2027 target

Equinix’s own materials describe the current status in three different ways, and the differences are not trivial.

The press release states the offering will be available starting in Q1 2027. The company’s product release note describes it as announced in preview, adding that further details on availability, deployment locations and customer access will be shared in future updates. The Horizon recap on the Equinix blog says it launches in Q1 2027.

Whether Q1 2027 represents general availability, limited availability, or simply the first customer deployments is therefore an open question. Equinix has not published a signup page, waitlist or preview enrolment path, so there is currently no mechanism for a customer to access the offering at all. The press release also carries a standard forward-looking-statements disclaimer, which applies to the date.

What isn’t public yet

For an enterprise evaluating infrastructure, the list of undisclosed items is longer than the list of disclosed ones.

DisclosedNot disclosed
Announcement date and venueInitial deployment metros or countries
The three partner companies and their rolesPricing model or indicative cost
Multitenant and dedicated single-tenant modesGPU models, node counts or rack configurations
Three target use casesCapacity figures
Stated Q1 2027 availabilityService level agreements
Together AI’s 200+ open-source model platformPerformance benchmarks or latency measurements
Delivery via Equinix FabricNamed customers, pilots or design partners

Trade publication eWeek reached the same conclusion, reporting that Equinix withheld its rollout locations, its pricing, the specifics of its hardware build, how much capacity it will offer and what performance it will commit to — and arguing that buyers need all of that before they can weigh the offering against production-inference options already on the market.

No technical whitepaper or architecture document for Inference Exchange has been published. No standalone press release from NVIDIA or Together AI accompanied the announcement; both companies are represented by executive quotes within Equinix’s release.

Absence of disclosure is not evidence of absence. These details may simply not be finalised. But they are what a buyer would need.

The footprint is not the rollout

The announcement is dense with impressive numbers — more than 280 data centres across 77 metros, 230 cloud on-ramps, and over 10,500 interconnected businesses, with eight of the top ten AI model providers and nine of the top ten AI clouds deployed with Equinix. The company’s Horizon recap adds that its metro footprint puts inference within roughly 10 milliseconds of the majority of internet users worldwide, across 77 metros in 36 countries.

Every one of those figures describes Equinix as a company. None describes where Inference Exchange will run.

That distinction is easy to lose and consequential to get wrong. eWeek drew the comparison explicitly, noting that Equinix’s September materials make no site-level deployment commitment equivalent to the one attached to its earlier Distributed AI Hub announcement in March. Capacity Equinix is adding elsewhere should not be assumed to serve Inference Exchange until launch metros are named.

The 10-millisecond figure deserves particular care: it is a measure of network proximity across Equinix’s footprint, not a measured inference result. Real inference latency also depends on model loading, scheduling, queueing, accelerator selection, batch size and the inference software itself. Equinix’s claim that the offering will cut time-to-first-token is directional and unquantified.

Who bills the customer

The single most operationally useful detail to emerge from the announcement did not appear in any Equinix document.

CNBC reported from the event that Together AI will be the seller of record, billing end customers who use the programme, and that financing details of the arrangement were not disclosed.

This is single-source reporting of a verbal statement, and it should be treated as reported rather than confirmed. If accurate, though, it has real consequences: contracting, invoicing and support escalation would run through Together AI, with Equinix providing the facilities and interconnection beneath. That is a meaningfully different commercial relationship from buying colocation or connectivity from Equinix directly, and it is worth verifying before any procurement conversation.

Where Fabric One fits

Equinix Fabric One was announced the same day, and the two are easy to conflate. They are separate products on separate schedules.

Fabric One is a managed connectivity service; Equinix says it enters beta later in 2026, with general availability planned for 2027 beginning in North America. No dependency between Fabric One and Inference Exchange has been stated in either direction.

What this means for enterprise buyers

Nothing about Inference Exchange is procurable today. Evaluation is limited to conversations with vendors.

More usefully, the offering cannot yet be compared against its alternatives on the terms that matter. Without pricing, a cost comparison against hyperscaler inference APIs, specialist GPU clouds or self-managed NVIDIA clusters is not possible. Without launch metros, the latency and data-residency arguments cannot be tested against an organisation’s actual geography. Without hardware configurations or benchmarks, throughput and cost-per-token projections are guesswork.

Three disclosures would change that: launch metros, a pricing model, and hardware configuration. Until they arrive, Inference Exchange belongs on a watch list rather than in a budget.

What’s changed since the announcement

As of 18 September 2026, the only material follow-up has been Equinix’s own Horizon recap, published 5 September, which restated the Q1 2027 timing and added the footprint proximity framing. No launch metros, pricing, hardware detail, customer names or schedule changes have been announced in the intervening two weeks.

DateDevelopment
2 September 2026Equinix Inference Exchange announced at Equinix Horizon, alongside Equinix Fabric One
5 September 2026Equinix publishes Horizon recap, restating Q1 2027 timing and footprint proximity figures
18 September 2026No further disclosures on metros, pricing, hardware or customers as of this writing
Q1 2027Stated availability target for Equinix Inference Exchange
Late 2026 → 2027Fabric One (separate product): beta later in 2026, general availability in 2027 starting in North America

Bottom line

Equinix has announced a coherent architecture with credible partners and a specific date. The logic behind it — that where inference runs is now a strategic decision rather than an infrastructure afterthought — is sound, and Equinix’s interconnection density is a genuine asset in that argument.

But an architecture is not a product. Until Equinix names launch metros, publishes pricing, and specifies hardware, Inference Exchange remains a statement of intent with a delivery date attached. The roughly two quarters between now and Q1 2027 are the window in which that changes — or doesn’t.

Frequently asked questions

What is Equinix Inference Exchange? A distributed AI inference offering announced by Equinix on 2 September 2026, combining Equinix data centres and Fabric connectivity, NVIDIA Enterprise Reference Architectures, and Together AI’s inference platform with access to more than 200 open-source models.

Is Equinix Inference Exchange available now? No. Equinix’s product release note describes it as announced in preview, with availability, deployment locations and customer access details to follow. There is no published signup or waitlist.

When will Equinix Inference Exchange launch? Equinix states availability begins in Q1 2027. Whether that means general availability or a more limited first release has not been clarified — the company has used several different formulations across its own materials.

Which locations will Equinix Inference Exchange run in? Not publicly disclosed. Equinix’s published figures — 280+ data centres across 77 metros — describe the company’s overall footprint, not the offering’s deployment map.

Which NVIDIA hardware does Equinix Inference Exchange use? Not publicly disclosed. Equinix has named NVIDIA’s Enterprise Reference Architecture programme but has not specified GPU models, node configurations or which reference design will be deployed.

How much does Equinix Inference Exchange cost, and who bills the customer? No pricing has been published. CNBC reported from the announcement that Together AI will act as seller of record and bill end customers, though this has not been confirmed in any Equinix document.

How is Equinix Inference Exchange different from Equinix Fabric One? Fabric One is a separate managed connectivity service announced the same day, with beta planned for later in 2026 and general availability in 2027 starting in North America. Equinix has not stated any dependency between the two.

Was this article helpful?

Rate this article from 1 to 5 stars.

Average rating: 0 / 5. Reader ratings: 0

No ratings yet. Be the first to rate this article.

Thank you for your feedback

Help us improve this article.

What information was missing or could be improved?

RELATED ARTICLES
Advertisement

Most Popular