Site and power readiness
The power path, the pad, and the access route, confirmed before anything is built. Stage 01.
INSIGHTDeployment · DP-06 · Stage 06 of 06
Operations and maintenance is the stage that lasts. It runs six workstreams — monitoring and alarm response, preventive maintenance, corrective maintenance and spares, access control, upgrade planning, and end-of-life planning — against a machine watched remotely and touched on a schedule. This guide covers what sets the intervals and where a modular unit changes the answer.
PUBLISHED LAST VERIFIED BY JOSEF ELIMELECHREVIEWED PODOS AI ENGINEERING
What stage 06 delivers
Monitoring and alarm response runs without pause. Preventive maintenance, corrective work and spares, access control, upgrade planning, and end-of-life planning run on triggers.
Six subsystems, each with an owner, a trigger, and a recorded result. Intervals come from the equipment manufacturer and the governing code.
Three tiers — on-unit, regional pool, factory or vendor — sorted by whether holding a part costs less than the downtime its lead time would cause.
Loop headroom, power train, fabric, and physical access, each cleared before a hardware refresh becomes a purchase decision.
The deployment chain
Stages 01 through 05 each end on a fixed date. Stage 06 inherits every decision they made.
The power path, the pad, and the access route, confirmed before anything is built. Stage 01.
Where the loop, the power train, and the fabric are fixed for the life of the unit. Stage 02.
Integration and test happen in the plant rather than on the site. Stage 03.
The unit moves by road to the prepared pad and is set in place. Stage 04.
Where the baseline every later trend is measured against gets set. Stage 05.
You are here. The first five stages end on a fixed date; this one does not.
Stage scope
The first five stages of a modular deployment end on a fixed date. Operations does not. Every earlier decision — the loop fixed in configuration engineering, the baseline set at commissioning — becomes a routine or a recurring problem here, one-way. Uptime Institute's 2025 survey of 800+ operators still finds outages common and staffing a persistent constraint: reasons to design O&M before the unit ships.[1]
A maintainable unit is one where every serviceable part can be reached, isolated, and replaced by one technician without shutting down the load. Four requirements follow. Clearance: room to withdraw a pump, a filter cartridge, or a battery string. Isolation: valves and breakers placed so one subsystem drops out while the rest runs, with obvious lockout points. Fluid handling: dripless disconnects, a drain and fill path, and a rehearsed procedure for breaking a coupling on a live loop.[3] Environmental separation: an outdoor-rated enclosure opens into weather, so ingress rating and corrosion class decide what can be serviced in the rain — and enclosure type ratings and IP codes do not translate cleanly both ways.[9]
In a factory-built unit all four are design parameters tested before shipment rather than negotiated with whatever the room turned out to be — but the envelope cannot be widened later. Clearance that was wrong in the factory is wrong on site.
Preventive maintenance
Intervals belong to the equipment manufacturer and the governing code, not to a generic table. What is generic is the register: the subsystems that must each have an owner, a trigger, and a recorded result.
| Code | Subsystem | Task | Consequence of skipping |
|---|---|---|---|
| OM-01 | Coolant chemistry | Sample and correct inhibitor level, conductivity, and biological growth.[3] | Corrosion products reach the cold plates and raise die temperature before an alarm fires. |
| OM-02 | Filtration | Change on differential pressure, not on a calendar date.[3] | A blinded filter starves rack flow; a bypassed one sends particulate to the cold plates. |
| OM-03 | Pumps, CDU, disconnects | Rotate redundant pumps, verify failover, inspect seals, rehearse isolation.[4] | Redundancy that is never exercised is an assumption, not a capability. |
| OM-04 | Heat rejection | Clean coil surfaces; verify approach temperature against the design ambient.[2] | Approach drifts with fouling, so capacity is lost on the hottest days. |
| OM-05 | Power train | Thermography on terminations, breaker and trip-unit testing, battery and generator load tests.[10] | Hot terminations and untested batteries are the classic outage. |
| OM-06 | Life safety and access | Test detection, alarm, and suppression per code; review access authorizations.[7] | Drifted detection is worse than none, because it is trusted. |
Reading the register
Loop work (OM-01, OM-02) is the habit air-cooled operators never had to build; federal-lab guidance treats it as core practice, not an add-on.[4] Power-train work (OM-05) has the shortest path to a full outage, which is why reliability practice counts maintenance inside the reliability calculation.[10] Access records are controls only while somebody audits them.[8]
Telemetry earns its cost when it changes what a technician does. Server health, thermal, and power sensors are readable out-of-band through a vendor-neutral model, so IT telemetry survives an unresponsive operating system.[5] Facility telemetry adds the loop — supply and return temperature, differential pressure, flow, filter delta-P — and the electrical side, where Class-A methods define how dips, unbalance, and harmonics are measured rather than sampled.[6] The instrumentation architecture behind that is covered in the monitoring and controls guide.
The operational point is narrower: trends set intervals. Filter changes follow differential pressure, coil cleaning follows approach temperature, pump service follows vibration.
Spares strategy
One rule sorts the whole parts list: hold a part locally when holding it costs less than the downtime its lead time would cause. Three tiers fall out.
Anything whose failure stops compute and whose replacement is a hand-tool job: filter cartridges, coolant charge, fans, seals, optics, fuses, a few drives. These live in the unit, not in a supply chain.
Pumps, CDU modules, power modules, breakers, battery strings, spare nodes. Shared across units, sized against failure rates, reachable inside a service-level window.
Long-lead assemblies needing requalification: transformers, switchgear, custom manifolds. Managed by contract lead times, not inventory.
Standardization is what makes this affordable. When every unit is the same build, one pool covers the fleet and every technician has seen the layout. Individually engineered rooms can pool nothing — the hidden operating cost of bespoke construction.
A calendar-only plan services healthy equipment and misses degrading equipment.
06
Subsystems needing an owner, a trigger, and a recorded result
Upgrades and lifecycle
Compute turns over faster than the infrastructure holding it, so a lifecycle plan is a plan for repeated re-population. Each refresh clears four gates before it becomes a purchase decision.
| Gate | Question | What it constrains |
|---|---|---|
| 01 | Does the loop have headroom at the same supply temperature? | Cold plates, flow rate, and whether heat rejection must be resized. |
| 02 | Does the power train carry the new draw with redundancy intact? | Distribution capacity, protection settings, battery autonomy. |
| 03 | Does the fabric match the new node count and port speeds? | Optics, cabling, and when redesign beats adaptation. |
| 04 | Is the physical work possible in the window, with the access you have? | A rolling in-place swap versus a unit taken out of service. |
In the product
Density is the variable that moves: the same 2025 survey shows fleet rack densities rising into the 10–30 kW band, so plan for the next generation asking for more.[1] A lifecycle plan also needs an ending — decommissioning, coolant recovery, data destruction, and, for a relocatable unit, moving it to where demand went. Those requirements sit with enclosure design and safety and security.
PODOS treats the maintenance register as part of the product definition, not a document written after handover. Because each PODOS Pod is designed as a standardized 1-MW building block and designed for 128 GPUs, service procedures, spares lists, and telemetry maps are identical across units — which is what makes a shared pool possible. The closed-loop direct-to-chip cooling system and the power architecture are commissioned in the factory, so the site inherits a known baseline to trend against.
Growth repeats stages 01 through 06 for another unit rather than re-opening a construction program, and PODOS targets a 90-day window from order to commissioning for a standard unit. Terms here are defined in the AI infrastructure glossary, and the buy-versus-rent framing is in the on-prem versus cloud comparison.
HONEST LIMITS
A remotely monitored, scheduled-touch unit is not universally the right answer.
QUESTIONS
Six workstreams: monitoring and alarm response, preventive maintenance on the cooling and power plant, corrective maintenance and spares, access control, upgrade planning, and end-of-life or relocation planning. Only the first is continuous.
It adds a discipline rather than a difficulty. A liquid loop brings coolant chemistry, filtration, and leak-detection tasks an air-cooled room does not have, and technicians break fluid couplings instead of sliding servers out of airflow. In exchange, the thermal envelope is far more stable.
Fewer than a conventional room of equivalent capacity, but never zero. Remote telemetry handles observation and first-line diagnosis; scheduled work, component replacement, and anything involving fluids or energized equipment needs a technician on site.
Bring the response window, the spares posture, and the refresh horizon. The configurator walks the same variables an engineering review would.