For a GPUaaS cloud, choose interconnects by mapping every link in the cluster to a reach class first, then applying three constraints: power per rack, consistency of latency under load, and supply continuity across build phases. Copper covers intra-rack and adjacent-rack links at the lowest power. Multimode optics cover in-row. Single-mode covers the rest. Utilisation economics make reliability worth more than unit price.
In a general-purpose cloud, the network is overhead. In a GPU cloud, it is part of the product. Customers rent capacity by the hour, and a fabric that stalls a distributed training job is directly consuming revenue.
That changes the selection criteria. Consistency matters more than peak performance. A link that is fast most of the time and occasionally retrains is worse than one that is marginally slower and never varies, because the slowest link sets the pace of a collective operation.
GPU clusters have a small number of repeating link types. Getting the map right first prevents the common failure of standardising on one technology and then using it everywhere, including where it is a poor fit. Note that distance sets what is possible, not what is mandatory. Short links can be run on optics rather than cable, and in many AI builds they are, for reasons covered below the table.
| Link | Typical distance | Copper option | Optical option |
| GPU server to top of rack, same cabinet | Under 3 m | Passive DAC | SR or DR pluggable optics, or AOC |
| Server to switch, tall or shared cabinet | 3 to 5 m | ACC | SR or DR pluggable optics, or AOC |
| Adjacent rack, row-level | 5 to 10 m | AEC | DR pluggable optics, or AOC |
| Leaf to spine, in-row | Up to 100 m | Not applicable | DR8 single-mode, increasingly the default. SR8 multimode or AOC where multimode plant exists |
| Hall to hall | Up to 500 m | Not applicable | DR8 single-mode |
| Building or campus | Up to 2 km | Not applicable | 2xFR4 single-mode |
Leaf to spine has moved decisively towards single-mode. Large fabrics increasingly run DR-variant optics at 100G, 400G and 800G rather than multimode, because a single fibre type across the estate simplifies cabling, patching and spares, and the plant carries the next speed step without being replaced. Multimode remains valid where it is already installed and in enterprise-scale estates, but it is no longer the default at AI or hyperscale.
The top three rows are a genuine choice rather than a rule. Copper wins on power, cost and latency within its reach. Optics win where a link has to be serviceable without disturbing the cable plant, where cable bulk behind the switch is restricting airflow, or where an operator wants one connectivity model across the whole cluster rather than four. Deciding that deliberately, per link class, is the point of the exercise. ATOP builds every option in that table, in copper and optics, at 400G and 800G, which means the map can be filled in link by link rather than compromised to fit a narrow catalogue.
GPU racks are already at the limit of what most facilities can power and cool. Every watt spent on the network is a watt not spent on compute, and in a colocation environment it may be a watt that is simply not available.
This is the strongest argument for using copper wherever reach allows. A passive DAC adds essentially nothing to the power budget and no meaningful latency. Active copper and AEC add a small amount in exchange for reach. Optical modules add the most, which is entirely reasonable when the distance requires it and wasteful when it does not.
Where optics are required, module architecture matters. Silicon photonics and linear pluggable designs reduce power by simplifying the electrical path. ATOP's 800G DR8 LPO silicon photonics variant is built for exactly this constraint.
Distributed training synchronises. A collective operation completes when its slowest participant completes, so the tail of the latency distribution determines job time, not the mean.
Three things push the tail out. Link flaps that force retraining. Marginal signal integrity that drives FEC correction up and occasionally past the correction limit. And thermal throttling in modules running close to their limit in a hot rack.
All three are avoidable through selection. Choose the shortest technology that covers the reach, since every conversion and retiming step adds a place for variance to enter. Specify operating temperature range against the real rack environment rather than the room. And qualify on measured error counters under load rather than on a link light.
In a single-tenant cluster, a bad link is an internal problem. In GPUaaS, it is a customer-facing incident with an SLA attached, and the customer has no visibility into the cause.
Field replacement is also harder. A rack running paying workloads cannot be taken down for module swaps at convenience. This pushes the value of consistency well above unit price, because the cost of a failure includes lost rental revenue, engineering time and reputational damage that does not appear in any procurement comparison.
GPU clouds grow in increments as capacity sells. The modules qualified in the first phase need to be available in the same specification for later phases, often more than a year apart.
That is a manufacturing question. Products built to a fixed design in owned facilities are the same batch after batch. ATOP designs, builds and tests in its own facilities and holds full bills of materials with component-level traceability, so a specification can be locked for the life of a build programme.
ATOP also sells direct, which removes a layer from both the technical conversation and the lead time. Fulfilment hubs in Denmark, the United States and Singapore support most build regions.