INSIGHTEngineering · Density

High-density GPU infrastructure, rack by rack

High-density GPU infrastructure is the practice of concentrating accelerated compute into racks whose sustained load exceeds what conventional room air cooling can serve — and then redesigning power delivery, heat removal, network fabric, clearances, and floor loading around that concentration. Density is not a specification you buy; it is the constraint set that every other decision in the room has to satisfy at once.

PUBLISHED LAST VERIFIED BY JOSEF ELIMELECHREVIEWED PODOS AI ENGINEERING

5
Density design regimes
10
Site readiness checks
11
Standards and sources cited

What you need to know

01

Density is a constraint set, not a spec

Concentrating accelerated compute redraws power delivery, heat removal, fabric, clearance, and floor loading at the same time.

02

The hardware moved first

Vendors now ship rack-scale accelerator systems liquid-cooled, while fleet-wide densities are still climbing into the 10–30 kW band.[1]

03

Cooling method follows density

Once the sustained load is fixed the cooling approach is implied rather than chosen — and every liquid design keeps a residual air path.

04

Weight stops more projects than heat

The governing number is usually the point load under a caster, not the distributed floor rating — and the delivery route gets checked last.

The pull and the push

Why density became the design driver

Training and inference clusters are latency-sensitive in a way general enterprise workloads are not. Accelerators that must exchange gradients or activations every few milliseconds pay for every metre of cable between them, so the fastest way to build a coherent machine is to put more of it in less space. That is the pull. The push is hardware: vendors now ship rack-scale accelerator systems as single liquid-cooled units — NVIDIA's GB200 NVL72, for example, places 72 GPUs and 36 CPUs in one rack acting as a single NVLink domain, and ships it liquid-cooled rather than offering an air variant.[4]

The installed base is moving more slowly than the hardware. Uptime Institute's 2025 survey of more than 800 operators shows rack densities rising into the 10–30 kW band fleet-wide[1]— well below what a modern accelerator rack demands. That gap is the whole problem: most existing rooms were designed for an airflow regime that the newest hardware has already left, and ASHRAE's TC 9.9 committee published a dedicated white paper on why liquid cooling is expanding into mainstream facilities for exactly this reason.[3]

Electrical

Power delivery per rack

At conventional density a rack is fed by whips from a floor PDU and nobody thinks hard about it. At high density the rack becomes a small electrical room. Three practical things change. First, the feed moves from cord-and-receptacle distribution toward overhead busway or dedicated panel feeds, because the conductor count and ampacity stop fitting under a floor. Second, phase balance becomes a per-rack concern rather than a per-row one: an accelerator rack draws a flat, high, sustained load, so imbalance shows up as heat and neutral current instead of averaging out. Third, overcurrent devices and conductors must be sized for continuous operation under the National Electrical Code, not for a nameplate peak the equipment never returns from.[9]

Redundancy has to be decided explicitly, because AI workloads split into two very different populations. Control planes, storage, and network fabric behave like classic critical load and justify concurrent-maintainable topology. Training jobs are checkpointed and restartable, so some operators accept a single path to the compute rack and spend the capital on capacity instead. Either choice is defensible; what is not defensible is assuming rack-level redundancy that the upstream distribution does not actually provide — the class of error the IEEE 3006 reliability-analysis series exists to surface.[11] The upstream chain that feeds all of this is covered in data center power architecture.

The ladder

What changes as density climbs

Density is best read as a sequence of design regimes rather than a number. Each step changes the cooling method, and each cooling method drags a different constraint to the front of the review.

#RegimeTypical cooling approachConstraint that dominates
HD-01Conventional enterprise racksRoom or in-row air handling with contained aisles; standard single- or dual-corded rack PDUs.Floor area and aisle airflow, not the rack itself. Structural and clearance rules are the legacy ones.
HD-02Dense air, containedContainment becomes mandatory; blanking panels, aisle pressure control, and ASHRAE class discipline decide whether the room holds.Fan energy and inlet temperature spread across the rack face; hot spots at the top of the rack.
HD-03Air-assist transitionRear-door heat exchangers or in-row liquid-to-air units move heat into water without touching the server internals.Facility water availability and floor space for the exchanger; a retrofit path more than a destination.
HD-04Direct-to-chip liquidCold plates on GPUs and CPUs, rack manifolds, quick disconnects, and a CDU; a residual air path stays for everything the plates do not touch.Coolant chemistry and flow balancing become operations disciplines; leak detection and isolation are engineered, not optional.
HD-05Rack-scale accelerator systemsVendor-integrated liquid-cooled racks that behave as one machine rather than a shelf of servers.The rack is now a single procurement, power, cooling, and service unit — partial population and mixed vendors get harder.

Regime boundaries are engineering conventions, not thresholds published by a standards body. The only fleet-wide measurement referenced above is Uptime Institute's 10–30 kW density band.[1]

Thermal

Matching cooling to the load

The cooling method is not a free choice once density is fixed — it is implied by it. ASHRAE's thermal guidelines define the air classes IT vendors design to, along with a high-density air class and facility water classes for liquid-cooled equipment; the class the hardware accepts sets the supply temperature the facility has to deliver.[2] Below the air ceiling, the work is containment discipline: blanking panels, aisle pressure, and inlet temperature uniformity across the full rack face. Above it, heat has to move into liquid, either at the rack door with a rear-door heat exchanger or at the die with cold plates. The Open Compute Project maintains vendor-neutral requirements across that whole range — cold plates, CDUs, rear-door exchangers, and heat reuse — so a loop can serve hardware from more than one vendor.[5][6] Its accelerator-infrastructure guidelines address multi-GPU systems specifically.[7]

The detail most often missed at design time is the residual air fraction. Cold plates capture heat only from the components they contact; voltage regulators, drives, NICs, and power supplies still reject heat to air. A high-density room therefore runs two cooling systems, and the smaller one still has to be engineered rather than inherited. Federal-lab guidance on liquid cooling in the rack treats this dual path — and the retrofit sequencing it implies — as standard practice.[8] The loop itself is covered in detail in direct-to-chip liquid cooling.

Physical envelope

Clearances, weight, and floor loading

High density compresses compute into a smaller footprint but does not compress the space around it. Service envelopes grow rather than shrink: a liquid-cooled rack needs rear access for manifolds and couplings, front access for chassis service, room for door swing, and the working space electrical code requires in front of equipment likely to be examined while energized.[9] The overhead zone is equally contested — busway, cable tray, piping, and fire protection all want the same volume, and NFPA 75 governs the fire-protection design of the IT equipment area itself.[10] A layout that satisfies the floor plan and fails the overhead section is a common and expensive discovery.

Weight is the constraint most often checked last and most likely to stop a project. Accelerator chassis, busway, manifolds, and the coolant charge concentrate mass into a footprint no larger than a legacy rack, and the governing figure is usually not the distributed floor rating but the point load under each caster or leveling foot. Raised floors have both a rated distributed load and a rated point load, and the second is what fails first. Two checks belong in every high-density design review: the structural capacity of the specific slab or raised floor against the manufacturer's published weights, and the delivery route — dock height, door widths, corridor turns, elevator capacity, ramp angles — from the truck to the final position. Seismic anchoring requirements, where local code imposes them, are a third.

Cold plates capture heat only from the components they contact — so a high-density room always runs two cooling systems, and the smaller one still has to be engineered rather than inherited.

PODOS AI Engineering · the residual air fraction

2

Cooling systems in a dense room

Interconnect

Network fabric at density

Three physical consequences follow, and all three are decided by the floor plan. Fabric layout, cooling layout, and service access are one design problem, and separating them is how rooms end up unserviceable.

NF-00

The traffic pattern changes first

Concentrating GPUs changes the traffic pattern before it changes anything else. Most of the bandwidth in an AI cluster is east-west — accelerator to accelerator — rather than north-south to users, so the fabric is designed around collective operations and tail latency, not aggregate throughput. Vendor rack-scale systems push this to its conclusion by making an entire rack one coherent interconnect domain.[4]

NF-01

Reach

The distance between compute racks and the switching tier sets whether links land on direct-attach copper, active copper, or optics, and optics are both a cost and a failure population.

NF-02

Cable mass and bend radius

High-radix fabrics put a very large number of terminations at the back of the rack, exactly where liquid manifolds and quick disconnects also live.

NF-03

Airflow

A dense rear cable bundle in an air-assisted regime is a thermal obstruction, not just an aesthetic one.

Before install

Site readiness checklist for a high-density rack

Ten items, in the order an engineering review tends to reach them. Each one has a specific failure mode when it is assumed instead of confirmed.

#DomainConfirm before installFailure mode if assumed
R-01PowerPer-rack circuit topology, phase balance across the rack face, and overcurrent sizing for continuous load.Breakers sized to nameplate rather than to code continuous-load rules trip under sustained training runs.
R-02PowerRedundancy intent: dual feed to the rack, single feed with restartable jobs, or something in between.Redundancy assumed at the rack but absent upstream — the failure mode IEEE 3006 reliability analysis exists to expose.
R-03CoolingCooling capacity matched to the actual sustained rack load, with the residual air fraction sized separately.Cold plates carry the GPUs while regulators, drives, and PSUs overheat in an under-sized air path.
R-04CoolingFacility water supply temperature and flow against the class the IT hardware accepts.A loop too cold risks condensation; too warm and the plates cannot hold junction temperature at full load.
R-05FabricCable plant sized for the east-west topology, including bend radius, tray fill, and optics reach.Cable bulk blocks the rear service access the liquid manifolds need.
R-06FabricPlacement of the switching tier relative to the compute racks and the reach budget that implies.Topology decided after the floor plan, forcing longer, more expensive optics or an extra hop.
R-07ClearanceFront and rear service envelopes, door swing, and the working space electrical code requires in front of energized gear.A rack that fits the floor plan but cannot legally or physically be serviced in place.
R-08ClearanceOverhead zone allocation between busway, cable tray, piping, and fire protection.Two trades designing into the same overhead volume; discovered during installation, not design.
R-09StructureDistributed floor load and the point load under casters or leveling feet, against the rated structure.A slab or raised floor that passes on average and fails under a single foot.
R-10StructureThe delivery path: dock height, door widths, corridor turns, elevator capacity, and ramp angles.Equipment that clears the room but not the route to it.

HONEST LIMITS

When high density is not the right answer

Density is a means, not a goal. There are cases where the honest recommendation is to build wider rather than denser.

  • The workload does not need adjacency. Density earns its cost when jobs span many GPUs and depend on low-latency interconnect. Embarrassingly parallel or storage-bound work runs perfectly well at conventional density and cheaper.
  • The building cannot take the load. Older shells with limited structural capacity, no facility water, or a fixed electrical service will consume more in retrofit than the density saves.
  • The operations team has no liquid experience. A liquid loop is an industrial system with chemistry, filtration, and leak procedures. Without trained staff or a service contract, density transfers risk from the design to the operator.
  • Growth is uncertain. Concentrating capacity into a small number of very large racks makes the increment coarse: the next unit of capacity is a whole rack, not a shelf.
  • Availability targets exceed the design. Uptime Institute's survey work continues to find outages a routine industry experience; density concentrates the consequence of a single rack-level fault, so the availability strategy has to be decided before the density is.

In the product

How PODOS handles the density constraint set

The list above is long because, in a conventional build, each item is resolved by a different party at a different time. PODOS moves the resolution into the factory. Each PODOS Pod is designed as a standardized 1 MW building block and designed for 128 GPUs, with power distribution, closed-loop liquid cooling, rack structure, and network paths engineered together as one enclosure rather than negotiated across trades on a site. Because the structural, clearance, and cooling relationships are fixed and tested before shipment, the site work reduces to the interfaces — service, water or heat rejection, and fibre — which is why PODOS targets a 90-day window from order to commissioning for a standard unit.

The same reasoning runs through the rest of the platform, the deployment model, and the workloads these units are built for. For the build-versus-buy framing, see modular vs traditional AI data centers; terms used above are defined in the AI infrastructure glossary.

QUESTIONS

Frequently asked questions

What counts as a high-density rack?

There is no standards-body threshold. In practice a rack is treated as high density once its sustained load exceeds what conventional room air handling can serve at the aisle, which is why the term tracks cooling method more than a specific kilowatt number. Uptime Institute's 2025 survey of more than 800 operators shows fleet-wide densities climbing into the 10-30 kW band, while accelerator racks sold today ship liquid-cooled well above it.

Can a high-density GPU rack be air cooled?

Up to a point, and the point is economic rather than physical. Air needs large temperature differences and high fan power to move meaningful energy, so as density rises the fan energy, aisle airflow, and floor area needed per rack all grow faster than the compute does. Vendors of the densest accelerator racks now ship them liquid-cooled rather than offering an air-cooled variant.

How much does a fully populated liquid-cooled rack weigh?

Enough that it must be checked against the structure, not assumed. Accelerator chassis, busway, manifolds, and the coolant charge all add mass in a footprint smaller than a legacy rack, so the governing number is usually the point load under the casters or leveling feet rather than the distributed floor rating. Confirm both against the manufacturer's published weight and a structural review of the specific slab or raised floor.

Does high density reduce total facility power?

No. Consolidating the same compute into fewer racks shortens cable runs and can cut fan and pump energy, but the silicon draws what it draws. Density changes where the heat appears and how efficiently it is removed; it does not change the IT load itself.

Bring your density target to engineering

Send the sustained rack load, the building constraints, and the growth increment. Engineering will tell you what a pod-based build looks like there.

Size your deploymentSee the deployment model