How to choose GPU interconnects for GPUaaS clouds

How GPUaaS providers should select transceivers and cables for GPU clusters: reach mapping, power per rack, tail latency, multi-tenant reliability and supply planning.

August 17, 2026

For a GPUaaS cloud, choose interconnects by mapping every link in the cluster to a reach class first, then applying three constraints: power per rack, consistency of latency under load, and supply continuity across build phases. Copper covers intra-rack and adjacent-rack links at the lowest power. Multimode optics cover in-row. Single-mode covers the rest. Utilisation economics make reliability worth more than unit price.

GPUaaS changes the economics of the network

In a general-purpose cloud, the network is overhead. In a GPU cloud, it is part of the product. Customers rent capacity by the hour, and a fabric that stalls a distributed training job is directly consuming revenue.

That changes the selection criteria. Consistency matters more than peak performance. A link that is fast most of the time and occasionally retrains is worse than one that is marginally slower and never varies, because the slowest link sets the pace of a collective operation.

Map every link class before choosing any product

GPU clusters have a small number of repeating link types. Getting the map right first prevents the common failure of standardising on one technology and then using it everywhere, including where it is a poor fit. Note that distance sets what is possible, not what is mandatory. Short links can be run on optics rather than cable, and in many AI builds they are, for reasons covered below the table.

Link Typical distance Copper option Optical option
GPU server to top of rack, same cabinet Under 3 m Passive DAC SR or DR pluggable optics, or AOC
Server to switch, tall or shared cabinet 3 to 5 m ACC SR or DR pluggable optics, or AOC
Adjacent rack, row-level 5 to 10 m AEC DR pluggable optics, or AOC
Leaf to spine, in-row Up to 100 m Not applicable DR8 single-mode, increasingly the default. SR8 multimode or AOC where multimode plant exists
Hall to hall Up to 500 m Not applicable DR8 single-mode
Building or campus Up to 2 km Not applicable 2xFR4 single-mode

Leaf to spine has moved decisively towards single-mode. Large fabrics increasingly run DR-variant optics at 100G, 400G and 800G rather than multimode, because a single fibre type across the estate simplifies cabling, patching and spares, and the plant carries the next speed step without being replaced. Multimode remains valid where it is already installed and in enterprise-scale estates, but it is no longer the default at AI or hyperscale.

The top three rows are a genuine choice rather than a rule. Copper wins on power, cost and latency within its reach. Optics win where a link has to be serviceable without disturbing the cable plant, where cable bulk behind the switch is restricting airflow, or where an operator wants one connectivity model across the whole cluster rather than four. Deciding that deliberately, per link class, is the point of the exercise. ATOP builds every option in that table, in copper and optics, at 400G and 800G, which means the map can be filled in link by link rather than compromised to fit a narrow catalogue.

 

Power per rack is the constraint that bites first

GPU racks are already at the limit of what most facilities can power and cool. Every watt spent on the network is a watt not spent on compute, and in a colocation environment it may be a watt that is simply not available.

This is the strongest argument for using copper wherever reach allows. A passive DAC adds essentially nothing to the power budget and no meaningful latency. Active copper and AEC add a small amount in exchange for reach. Optical modules add the most, which is entirely reasonable when the distance requires it and wasteful when it does not.

Where optics are required, module architecture matters. Silicon photonics and linear pluggable designs reduce power by simplifying the electrical path. ATOP's 800G DR8 LPO silicon photonics variant is built for exactly this constraint.

Tail latency, not average latency, is what customers feel

Distributed training synchronises. A collective operation completes when its slowest participant completes, so the tail of the latency distribution determines job time, not the mean.

Three things push the tail out. Link flaps that force retraining. Marginal signal integrity that drives FEC correction up and occasionally past the correction limit. And thermal throttling in modules running close to their limit in a hot rack.

All three are avoidable through selection. Choose the shortest technology that covers the reach, since every conversion and retiming step adds a place for variance to enter. Specify operating temperature range against the real rack environment rather than the room. And qualify on measured error counters under load rather than on a link light.

Multi-tenancy raises the cost of a fault

In a single-tenant cluster, a bad link is an internal problem. In GPUaaS, it is a customer-facing incident with an SLA attached, and the customer has no visibility into the cause.

Field replacement is also harder. A rack running paying workloads cannot be taken down for module swaps at convenience. This pushes the value of consistency well above unit price, because the cost of a failure includes lost rental revenue, engineering time and reputational damage that does not appear in any procurement comparison.

Build in phases, buy for the whole programme

GPU clouds grow in increments as capacity sells. The modules qualified in the first phase need to be available in the same specification for later phases, often more than a year apart.

That is a manufacturing question. Products built to a fixed design in owned facilities are the same batch after batch. ATOP designs, builds and tests in its own facilities and holds full bills of materials with component-level traceability, so a specification can be locked for the life of a build programme.

ATOP also sells direct, which removes a layer from both the technical conversation and the lead time. Fulfilment hubs in Denmark, the United States and Singapore support most build regions.

A selection checklist for GPUaaS

  • Produce a link map with measured distances for every repeating link class in the cluster
  • Use the shortest viable technology per class, defaulting to copper wherever reach allows
  • Model front-panel power per rack, not per module
  • Specify temperature range against real rack conditions under full load
  • Confirm host coding requirements for every switch platform in the design
  • Lock the specification and supply plan across all planned build phases before the first order

 

What interconnects are used in GPU clusters?

Neither is better in general. Copper uses less power, adds less latency and costs less within its reach, which is around 10 m at 400G and 800G. Fibre is required beyond that and is also chosen inside copper reach where links must be serviceable without disturbing the cable plant or where cable bulk is affecting airflow. Beyond the rack, single-mode DR optics have become the common choice at scale.

How does interconnect choice affect GPU cluster performance?

Distributed training synchronises across nodes, so job completion time is set by the slowest link rather than the average. Interconnect choices that reduce link instability, keep FEC correction well inside its budget and avoid thermal throttling reduce tail latency, which is what actually determines training throughput.

What should a GPUaaS provider ask an interconnect supplier?

Ask whether the supplier manufactures or sources, whether they own their firmware and can code for the switch platforms in use, what the measured power figure is at realistic case temperature, what documentation and traceability comes with each batch, and whether the same specification can be guaranteed across build phases twelve to eighteen months apart.
arrow-right