DP-02Deploy · Deployment stage

Configuration engineering: requirements become a specified unit

Configuration engineering is the stage that turns a workload requirement into a build specification. On a standardized modular unit it is a bounded selection problem, not a design project: workload profile, accelerator family, rack density, cooling selection, site interfaces, and operating model are each chosen from a fixed menu. The output is a configuration freeze — the signed specification the factory builds against and every later acceptance test references.

PUBLISHED LAST VERIFIED BY JOSEF ELIMELECHREVIEWED PODOS AI ENGINEERING

6
Configuration axes, in dependency order
4
Site interfaces to define
11
Lines in the freeze checklist

What stage 02 delivers

01

One deliverable: the freeze

The stage has one deliverable and one failure mode. The deliverable is the configuration freeze — the signed specification the factory builds against.

02

A bounded menu, not a design project

The architecture is already fixed, so configuration is the short step of choosing among options the platform already supports.

03

Order is the whole discipline

Six axes, each constraining the next. Working the chain backwards is the most common way a configuration fails to close.

04

One failure mode

Freezing a specification nobody traced end to end — an accelerator chosen without checking the density it implies.

The chain

Where stage 02 sits in the six-stage deployment model

Each stage inherits the decisions frozen in the one before it. This page is the stage-02 detail.

DP-01

Site and power readiness

The site-side half of the freeze checklist is worked here — service, interconnect, pad, and access route. Read stage 01.

DP-02 · THIS STAGE

Configuration engineering

You are here. Requirements become a signed build specification: six axes chosen in dependency order and frozen.

DP-03

Factory build and testing

The unit is built and tested against the frozen specification, indoors and off the critical path. Read stage 03.

DP-04

Transport and placement

The finished unit moves to site inside the transport envelope named in the freeze, and lands on its pad. Read stage 04.

DP-05

Commissioning

On-site acceptance testing, referenced back to the specification frozen in this stage. Read stage 05.

DP-06

Operations and maintenance

The operating model signed in the freeze becomes live monitoring, maintenance, and spares. Read stage 06.

Scope

What this stage actually decides

In a conventional build, configuration and design are the same activity, and they run for months because almost nothing is fixed. In a factory-built model the architecture is already fixed — enclosure, cooling topology, power distribution, rack geometry — so configuration is the short step of choosing among options the platform already supports. That is the whole trade: less freedom, far less schedule. This page is the stage-02 detail behind the six-stage deployment overview.

The stage has one deliverable and one failure mode. The deliverable is the configuration freeze. The failure mode is freezing a specification nobody traced end to end — an accelerator chosen without checking the density it implies, a density chosen without checking the cooling it forces, a cooling selection chosen without checking what the site can reject heat into.

Workload profile sets everything else

The first question is not which GPU. It is what the machine is for, and how hard it runs. A sustained training profile pins accelerators near their power ceiling for long stretches, so the thermal design is sized against a near-continuous load. An inference profile is spikier and more sensitive to latency and data residency, which often decides the site before it decides the hardware. A mixed profile must be sized against its worst sustained case — averaging is how thermal designs end up undersized.

Two assumptions belong in writing at this point: the sustained duty the cooling is sized against, and the horizon over which the unit must absorb a hardware generation or two. Both are what a later disagreement will be about.

Accelerator family and rack density

Choosing an accelerator family is no longer only a compute decision. Vendors now ship rack-scale packages with their own cooling and interconnect assumptions built in: NVIDIA's GB200 NVL72 is delivered as a liquid-cooled rack acting as a single NVLink domain, because an air-cooled build of that density is not on offer.[2] The interconnect topology travels with the choice too: NVLink Switch gives all-to-all GPU communication across the rack,[3] which is a different design question from the fabric between racks. Configuration has to accept both, or pick a different family.

Rack density then falls out of that choice, and it belongs in the freeze as a sustained figure across the refresh horizon rather than a day-one nameplate. Uptime Institute's 2025 global data center survey tracks the same shift across the operator base: fleet densities are climbing past what conventional air cooling handles comfortably.[1] The high-density GPU infrastructure page covers what that density does to the rest of the design.

Cooling selection follows density, not preference

Once density is fixed, the cooling selection is mostly determined. What remains open is the coolant class, the facility water supply temperature the installed hardware accepts, and the heat-rejection method at the boundary. ASHRAE's thermal guidelines name liquid-cooling facility water classes by their maximum supply temperature, and the choice is economic as much as thermal: the warmer the class the hardware tolerates, the more of the year heat can be rejected without compressors.[4]Open Compute's cold-plate requirements set the wetted-material and interface expectations that keep a multi-vendor loop stable across refreshes.[5]

The rejection method is the genuinely site-specific part, sized against climatic design conditions for the actual location — design dry bulb, extreme dew point, coincident wet bulb — not a national average.[9] A dry cooler consumes no water but needs a warmer loop or more surface area; evaporative rejection buys lower approach temperatures and pays in water. The loop itself is covered on the direct-to-chip liquid cooling page.

Dependency order

The six configuration axes, in dependency order

These are decided in order, because each constrains the next. Working the chain backwards — from a preferred cooling method or a preferred site — is the most common way a configuration fails to close.

AxisDecisionWhat is chosenDownstream consequence
CFG-01Workload profileTraining, inference, or mixed; sustained versus bursty duty; latency and data-residency limits.Sets every axis below it. A sustained training profile and a spiky inference profile produce different hardware from the same enclosure.
CFG-02Accelerator familyWhich accelerator platform is installed at integration, and how the vendor packages it at rack scale.Rack-scale vendor packages carry their own thermal and interconnect assumptions; accepting the family means accepting those.
CFG-03Rack densitySustained kW per rack across the refresh horizon, not the day-one figure.Decides whether air stays viable, how many racks the unit carries, and how much of the electrical budget is compute.
CFG-04Cooling selectionCoolant class, facility water supply temperature, and the heat-rejection method at the boundary.Warmer accepted supply water buys free cooling; the rejection choice sets the site water story and outdoor footprint.
CFG-05Site interfacesFour boundaries: electrical service, network handoff, heat rejection, physical placement.The only points where a standardized unit meets a non-standard world, so they hold most of the remaining risk.
CFG-06Operating modelWho monitors, who maintains, what telemetry leaves the unit, and to whom.Sets monitoring integration scope, spares, and response commitments. The last axis to stay flexible.

CFG-05

Site interfaces: where a standard unit meets a non-standard world

A standardized unit has exactly four boundaries, and configuring them is most of the remaining engineering.

ELECTRICAL

Service, ownership, harmonics

On the electrical side the specification states the service arrangement, transformer ownership, protection coordination, and who owns harmonic performance at the point of common coupling — IEEE 519 sets the limits rectifier-heavy loads are judged against.[6] The power architecture page describes what sits on the unit side of that boundary.

NETWORK

Demarcation, media, port speeds

On the network side, the freeze names the demarcation point, media type, port speeds, and cross-connect ownership. TIA-942 supplies the vocabulary for entrance rooms, pathways, and cross-connects that site agreements are written in,[7] while IEEE 802.3df defines the 400 and 800 Gb/s Ethernet layers now common on AI fabrics — and their very different copper, multimode, and single-mode reach classes, which is really a question about where the unit can sit relative to the meet-me point.[8] See networking and fiber for the fabric side.

THERMAL

Heat rejection and clearances

The thermal boundary is the heat-rejection equipment and its clearances.

PHYSICAL

Pad, anchorage, access route

The physical boundary is the pad, anchorage, access route, and the enclosure rating the site's exposure demands, for which the IP code is the usual shorthand.[10] The route is a hard constraint, not a formality: federal standards fix the National Network width at 102 inches and cap interstate gross vehicle weight at 80,000 pounds.[12]

Acceptance criteria

The configuration freeze checklist

A freeze is complete when every line below has a stated answer and a named owner. An open line is not a detail for later — it is a change order waiting to happen after procurement has started.

#What must be answered, and owned, before the freeze is signed
01Workload profile stated in writing, including the sustained duty the thermal design is sized against.
02Accelerator family and quantity fixed, with the vendor's thermal and interconnect requirements attached.
03Sustained kW per rack agreed for the refresh horizon, not only the first installed generation.
04Coolant class and facility water supply temperature selected and confirmed against the installed hardware.
05Heat-rejection method chosen, with the site's climatic design conditions named as the sizing basis.
06Electrical service arrangement fixed: voltage, transformer ownership, protection coordination, harmonic responsibility.
07Network handoff defined: demarcation point, media type, port speeds, cross-connect ownership.
08Physical placement resolved: pad, anchorage, clearances, access route, transport envelope.
09Enclosure rating stated for the site's exposure, corrosion, and altitude conditions.
10Operating model signed: monitoring scope, telemetry destination, maintenance responsibility, spares position.
11Change-order process agreed, naming which axes stay open and what reopening each one costs.

A freeze is complete when every line has a stated answer and a named owner. An open line is a change order waiting to happen after procurement has started.

PODOS AI Engineering · configuration freeze

11

Lines that must close before signature

In the product

How PODOS runs stage 02

The site-side half of this list is worked in stage 01; the data center readiness checklist covers it in detail, and unfamiliar terms are defined in the AI infrastructure glossary. Finally, the operating model decides what telemetry leaves the unit and in what form; Redfish is the vendor-neutral out-of-band model most monitoring and controls integrations are built against.[11]

PODOS keeps the architecture fixed so the menu can stay short. Each PODOS Pod is designed as a standardized 1 MW building block and designed for 128 GPUs, with the enclosure, closed-loop direct-to-chip cooling, and power distribution identical from unit to unit. Configuration therefore selects the accelerator family, the heat-rejection option, the service and network arrangements at the boundary, and the operating model — nothing structural. Holding the architecture constant is what makes the calendar predictable enough that PODOS targets a 90-day window from order to commissioning for a standard unit.

To work the axes above as a live selection, use the configurator. It walks the same order — profile, accelerators, density, cooling, site interfaces — and produces the outline a configuration freeze is written from.

HONEST LIMITS

When configuration engineering is not the right fit

A bounded configuration menu is an advantage only when the requirement fits inside it. It does not, in these cases:

  • The requirement is genuinely unknown. If the workload is still being discovered, freezing early gives a precise answer to the wrong question. Rent capacity, learn the load, then configure.
  • The design must leave the fixed architecture. A non-standard enclosure geometry, a bespoke rack pitch, or a cooling topology outside the platform's options is a design project and should be priced as one.
  • A single scale-up domain exceeds what one unit holds. When a model needs more tightly-coupled accelerators than a unit carries, the interconnect topology between units becomes the primary design problem.
  • The binding constraint is upstream. Configuration cannot shorten an interconnect queue, resolve a zoning dispute, or create water rights. If the site is the constraint, stage 01 owns the calendar.
  • The jurisdiction has no off-site construction path. Where the authority having jurisdiction will not accept factory inspection in lieu of site inspection, a site-built comparison deserves a fair hearing.

QUESTIONS

Frequently asked questions

What is configuration engineering in a data center deployment?

It is the stage that converts requirements into a build specification. On a standardized unit it is a bounded selection problem rather than a design project: workload profile, accelerator family, rack density, cooling option, site interfaces, and operating model are each chosen from a fixed menu and signed as a configuration freeze.

What is a configuration freeze?

The signed specification the factory builds against and every later acceptance test references. Once frozen, changes run as documented change orders, because procurement, fabrication sequencing, and factory test plans are all derived from it.

Which decision should be made first?

The workload profile. It sets the accelerator family, which sets rack density, which sets the cooling selection, which sets the heat-rejection and site-interface requirements. Working the chain in the other direction produces a specification that does not close.

Can the configuration change after the freeze?

Some of it, at a cost. Monitoring scope and operating-model choices stay flexible longest. Accelerator family, rack density, and anything touching the cooling loop or the electrical service become expensive once fabrication and long-lead procurement have started.

Work the six axes as a live selection

Start from the workload profile and let each axis constrain the next. The configurator produces the outline a configuration freeze is written from.

Size your deploymentSee the deployment model