Site and power readiness
The site-side half of the freeze checklist is worked here — service, interconnect, pad, and access route. Read stage 01.
DP-02Deploy · Deployment stage
Configuration engineering is the stage that turns a workload requirement into a build specification. On a standardized modular unit it is a bounded selection problem, not a design project: workload profile, accelerator family, rack density, cooling selection, site interfaces, and operating model are each chosen from a fixed menu. The output is a configuration freeze — the signed specification the factory builds against and every later acceptance test references.
PUBLISHED LAST VERIFIED BY JOSEF ELIMELECHREVIEWED PODOS AI ENGINEERING
What stage 02 delivers
The stage has one deliverable and one failure mode. The deliverable is the configuration freeze — the signed specification the factory builds against.
The architecture is already fixed, so configuration is the short step of choosing among options the platform already supports.
Six axes, each constraining the next. Working the chain backwards is the most common way a configuration fails to close.
Freezing a specification nobody traced end to end — an accelerator chosen without checking the density it implies.
The chain
Each stage inherits the decisions frozen in the one before it. This page is the stage-02 detail.
The site-side half of the freeze checklist is worked here — service, interconnect, pad, and access route. Read stage 01.
You are here. Requirements become a signed build specification: six axes chosen in dependency order and frozen.
The unit is built and tested against the frozen specification, indoors and off the critical path. Read stage 03.
The finished unit moves to site inside the transport envelope named in the freeze, and lands on its pad. Read stage 04.
On-site acceptance testing, referenced back to the specification frozen in this stage. Read stage 05.
The operating model signed in the freeze becomes live monitoring, maintenance, and spares. Read stage 06.
Scope
In a conventional build, configuration and design are the same activity, and they run for months because almost nothing is fixed. In a factory-built model the architecture is already fixed — enclosure, cooling topology, power distribution, rack geometry — so configuration is the short step of choosing among options the platform already supports. That is the whole trade: less freedom, far less schedule. This page is the stage-02 detail behind the six-stage deployment overview.
The stage has one deliverable and one failure mode. The deliverable is the configuration freeze. The failure mode is freezing a specification nobody traced end to end — an accelerator chosen without checking the density it implies, a density chosen without checking the cooling it forces, a cooling selection chosen without checking what the site can reject heat into.
The first question is not which GPU. It is what the machine is for, and how hard it runs. A sustained training profile pins accelerators near their power ceiling for long stretches, so the thermal design is sized against a near-continuous load. An inference profile is spikier and more sensitive to latency and data residency, which often decides the site before it decides the hardware. A mixed profile must be sized against its worst sustained case — averaging is how thermal designs end up undersized.
Two assumptions belong in writing at this point: the sustained duty the cooling is sized against, and the horizon over which the unit must absorb a hardware generation or two. Both are what a later disagreement will be about.
Choosing an accelerator family is no longer only a compute decision. Vendors now ship rack-scale packages with their own cooling and interconnect assumptions built in: NVIDIA's GB200 NVL72 is delivered as a liquid-cooled rack acting as a single NVLink domain, because an air-cooled build of that density is not on offer.[2] The interconnect topology travels with the choice too: NVLink Switch gives all-to-all GPU communication across the rack,[3] which is a different design question from the fabric between racks. Configuration has to accept both, or pick a different family.
Rack density then falls out of that choice, and it belongs in the freeze as a sustained figure across the refresh horizon rather than a day-one nameplate. Uptime Institute's 2025 global data center survey tracks the same shift across the operator base: fleet densities are climbing past what conventional air cooling handles comfortably.[1] The high-density GPU infrastructure page covers what that density does to the rest of the design.
Once density is fixed, the cooling selection is mostly determined. What remains open is the coolant class, the facility water supply temperature the installed hardware accepts, and the heat-rejection method at the boundary. ASHRAE's thermal guidelines name liquid-cooling facility water classes by their maximum supply temperature, and the choice is economic as much as thermal: the warmer the class the hardware tolerates, the more of the year heat can be rejected without compressors.[4]Open Compute's cold-plate requirements set the wetted-material and interface expectations that keep a multi-vendor loop stable across refreshes.[5]
The rejection method is the genuinely site-specific part, sized against climatic design conditions for the actual location — design dry bulb, extreme dew point, coincident wet bulb — not a national average.[9] A dry cooler consumes no water but needs a warmer loop or more surface area; evaporative rejection buys lower approach temperatures and pays in water. The loop itself is covered on the direct-to-chip liquid cooling page.
Dependency order
These are decided in order, because each constrains the next. Working the chain backwards — from a preferred cooling method or a preferred site — is the most common way a configuration fails to close.
| Axis | Decision | What is chosen | Downstream consequence |
|---|---|---|---|
| CFG-01 | Workload profile | Training, inference, or mixed; sustained versus bursty duty; latency and data-residency limits. | Sets every axis below it. A sustained training profile and a spiky inference profile produce different hardware from the same enclosure. |
| CFG-02 | Accelerator family | Which accelerator platform is installed at integration, and how the vendor packages it at rack scale. | Rack-scale vendor packages carry their own thermal and interconnect assumptions; accepting the family means accepting those. |
| CFG-03 | Rack density | Sustained kW per rack across the refresh horizon, not the day-one figure. | Decides whether air stays viable, how many racks the unit carries, and how much of the electrical budget is compute. |
| CFG-04 | Cooling selection | Coolant class, facility water supply temperature, and the heat-rejection method at the boundary. | Warmer accepted supply water buys free cooling; the rejection choice sets the site water story and outdoor footprint. |
| CFG-05 | Site interfaces | Four boundaries: electrical service, network handoff, heat rejection, physical placement. | The only points where a standardized unit meets a non-standard world, so they hold most of the remaining risk. |
| CFG-06 | Operating model | Who monitors, who maintains, what telemetry leaves the unit, and to whom. | Sets monitoring integration scope, spares, and response commitments. The last axis to stay flexible. |
CFG-05
A standardized unit has exactly four boundaries, and configuring them is most of the remaining engineering.
On the electrical side the specification states the service arrangement, transformer ownership, protection coordination, and who owns harmonic performance at the point of common coupling — IEEE 519 sets the limits rectifier-heavy loads are judged against.[6] The power architecture page describes what sits on the unit side of that boundary.
On the network side, the freeze names the demarcation point, media type, port speeds, and cross-connect ownership. TIA-942 supplies the vocabulary for entrance rooms, pathways, and cross-connects that site agreements are written in,[7] while IEEE 802.3df defines the 400 and 800 Gb/s Ethernet layers now common on AI fabrics — and their very different copper, multimode, and single-mode reach classes, which is really a question about where the unit can sit relative to the meet-me point.[8] See networking and fiber for the fabric side.
The thermal boundary is the heat-rejection equipment and its clearances.
The physical boundary is the pad, anchorage, access route, and the enclosure rating the site's exposure demands, for which the IP code is the usual shorthand.[10] The route is a hard constraint, not a formality: federal standards fix the National Network width at 102 inches and cap interstate gross vehicle weight at 80,000 pounds.[12]
Acceptance criteria
A freeze is complete when every line below has a stated answer and a named owner. An open line is not a detail for later — it is a change order waiting to happen after procurement has started.
| # | What must be answered, and owned, before the freeze is signed |
|---|---|
| 01 | Workload profile stated in writing, including the sustained duty the thermal design is sized against. |
| 02 | Accelerator family and quantity fixed, with the vendor's thermal and interconnect requirements attached. |
| 03 | Sustained kW per rack agreed for the refresh horizon, not only the first installed generation. |
| 04 | Coolant class and facility water supply temperature selected and confirmed against the installed hardware. |
| 05 | Heat-rejection method chosen, with the site's climatic design conditions named as the sizing basis. |
| 06 | Electrical service arrangement fixed: voltage, transformer ownership, protection coordination, harmonic responsibility. |
| 07 | Network handoff defined: demarcation point, media type, port speeds, cross-connect ownership. |
| 08 | Physical placement resolved: pad, anchorage, clearances, access route, transport envelope. |
| 09 | Enclosure rating stated for the site's exposure, corrosion, and altitude conditions. |
| 10 | Operating model signed: monitoring scope, telemetry destination, maintenance responsibility, spares position. |
| 11 | Change-order process agreed, naming which axes stay open and what reopening each one costs. |
A freeze is complete when every line has a stated answer and a named owner. An open line is a change order waiting to happen after procurement has started.
11
Lines that must close before signature
In the product
The site-side half of this list is worked in stage 01; the data center readiness checklist covers it in detail, and unfamiliar terms are defined in the AI infrastructure glossary. Finally, the operating model decides what telemetry leaves the unit and in what form; Redfish is the vendor-neutral out-of-band model most monitoring and controls integrations are built against.[11]
PODOS keeps the architecture fixed so the menu can stay short. Each PODOS Pod is designed as a standardized 1 MW building block and designed for 128 GPUs, with the enclosure, closed-loop direct-to-chip cooling, and power distribution identical from unit to unit. Configuration therefore selects the accelerator family, the heat-rejection option, the service and network arrangements at the boundary, and the operating model — nothing structural. Holding the architecture constant is what makes the calendar predictable enough that PODOS targets a 90-day window from order to commissioning for a standard unit.
To work the axes above as a live selection, use the configurator. It walks the same order — profile, accelerators, density, cooling, site interfaces — and produces the outline a configuration freeze is written from.
HONEST LIMITS
A bounded configuration menu is an advantage only when the requirement fits inside it. It does not, in these cases:
QUESTIONS
It is the stage that converts requirements into a build specification. On a standardized unit it is a bounded selection problem rather than a design project: workload profile, accelerator family, rack density, cooling option, site interfaces, and operating model are each chosen from a fixed menu and signed as a configuration freeze.
The signed specification the factory builds against and every later acceptance test references. Once frozen, changes run as documented change orders, because procurement, fabrication sequencing, and factory test plans are all derived from it.
The workload profile. It sets the accelerator family, which sets rack density, which sets the cooling selection, which sets the heat-rejection and site-interface requirements. Working the chain in the other direction produces a specification that does not close.
Some of it, at a cost. Monitoring scope and operating-model choices stay flexible longest. Accelerator family, rack density, and anything touching the cooling loop or the electrical service become expensive once fabrication and long-lead procurement have started.
Start from the workload profile and let each axis constrain the next. The configurator produces the outline a configuration freeze is written from.