UC-02Use case · Enterprise AI
When dedicated capacity earns its place
Three conditions have to hold at once: utilization is high and sustained, the data carries a residency requirement rented capacity cannot satisfy, and a real power path exists on land you control. Miss one and reserved cloud is the better answer.
- 3
- Conditions, all required
- 6
- Roles that must agree
- 10
- Checks before hardware
The test, in three parts
Utilization is sustained
Continuous fine-tuning, batch scoring, production inference. It runs most hours and grows in a straight line, not in bursts.
Residency is named, not preferred
A contract clause, a regulator's expectation, a jurisdiction rule. “We'd prefer in-house” does not survive a capital review.
A power path exists
Service headroom or an interconnection request with a credible energization date, on land the organization controls.
The economics
The workload profile that justifies owning capacity
Enterprise AI rarely looks like a research cluster. The load that justifies owned infrastructure is repetitive: continuous fine-tuning on proprietary corpora, batch scoring against internal records, and production inference behind an internal assistant or a customer-facing feature. It runs on a schedule, it runs most hours, and it grows in a straight line rather than in bursts.
That distinction decides the economics. An owned megawatt is paid for whether or not it is busy, so the comparison is never list price against list price — it is the fully loaded cost of the asset divided by the hours it actually runs. Experimental workloads with long idle stretches lose that comparison; steady, forecastable ones win it, which is why the underlying load is growing as it is. U.S. data centers consumed roughly 4.4% of national electricity in 2023, with Lawrence Berkeley National Laboratory projecting 6.7–12% by 2028,[1] and the IEA expects global data-centre electricity use to climb toward roughly 945 TWh by 2030, driven largely by AI.[2]
The second qualifier is data. A residency requirement is a requirement only when somebody can name it — a contract clause, a regulator's expectation, a jurisdiction rule. "We would prefer to keep it in-house" is a preference, and preferences do not survive a capital review. The two delivery models are compared directly in on-prem AI infrastructure vs cloud.
Decision map
Who decides, and what each role blocks on
On-site AI capacity requires the AI team, infrastructure, security, facilities, and finance to agree — plus two external parties nobody in the building controls. Projects stall on the role invited last.
| Code | Role | Owns | Blocks on |
|---|---|---|---|
| EA-R1 | Head of AI / ML platform | Model roadmap, utilization, scheduler policy | Whether today's silicon is still right two refresh cycles out. |
| EA-R2 | VP infrastructure / CIO | Capacity plan, operating model, remote hands | Who carries the pager for a facility IT has never owned. |
| EA-R3 | Security & data governance | Classification, access control, egress rules | Whether physical custody is a stated requirement or a preference. |
| EA-R4 | Facilities / real estate | Land, pad, loading, delivery route, heat rejection | Whether the site takes another megawatt without a service upgrade. |
| EA-R5 | Finance / procurement | Capital treatment, depreciation horizon, vendor risk | Utilization assumptions — the case collapses if the cluster idles. |
| EA-R6 | External: utility and AHJ | Service study, interconnection, permits, inspection | Energization date; acceptance of off-site-built equipment. |
Technical profile
A GPU cluster is not a bigger server room
The Uptime Institute's 2025 survey of more than 800 operators reports fleet rack densities rising into the 10–30 kW band,[3]while rack-scale AI systems ship liquid-cooled — NVIDIA's GB200 NVL72 puts 72 GPUs and 36 CPUs in one rack acting as a single NVLink domain, with no air-cooled equivalent on offer.[4] A room built for the first number does not quietly accommodate the second.
Power. Dense racks change the design upstream as much as the rack itself: service capacity, distribution topology, and the redundancy the workload actually needs — often lower for a training cluster than for a transactional system. Power architecture.
Cooling. ASHRAE's thermal guidelines define both air classes and facility water classes for liquid cooling,[5] and above a certain density the question stops being which class and becomes which loop. Direct-to-chip cooling.
Network. Multi-node training generally needs two fabrics: practice at scale separates the frontend network from a dedicated, non-blocking backend fabric carrying collective traffic between GPUs.[6] The IEEE standardized 800 Gb/s Ethernet in 2024,[7] so the components exist — the wide-area path back to the data estate is the part most plans underestimate. Networking and fiber.
Site survey
What the site has to provide
A modular unit removes the building program, not the site program. Six things still have to be true about the property before hardware selection means anything.
| Requirement | What “done” looks like | Why it bites |
|---|---|---|
| Power path | Service headroom, or an interconnection request with a credible energization date. | The usual reason an on-site plan slips from quarters to years. |
| Pad and access | A level pad with structural capacity and a legal road route for delivery. | Interstate limits are 80,000 lb gross and 102 in wide on the National Network; height is left to the states.[9] |
| Permitting | An AHJ that accepts factory-built equipment and plant-stage inspection. | Off-site construction standards define those roles, but adoption varies by jurisdiction. |
| Heat rejection | Somewhere for the heat to go: dry coolers, evaporative equipment, or a heat consumer. | Climate and water position decide it; evaporative permits are the slow path. |
| Network entrance | Fiber in, sized for operations traffic and data movement to the corporate estate. | A cluster with no path to its data is a stranded asset. |
| Operations coverage | Named remote hands, spares logistics, escalation path. | The organization now operates infrastructure instead of consuming a service. |
A private enclosure is a physical-control posture, not a compliance status. Compliance is produced by an audit against a named framework, by a qualified assessor, for a defined scope.
EA-05 · Control posture
Data residency and control: what a private site changes
Owning the enclosure changes three concrete things: where the data physically sits, who can put hands on the hardware, and who holds the keys to the management plane. Those map to an existing control vocabulary — NIST's SP 800-53 physical and environmental protection family covers access authorizations, physical access control and monitoring, visitor records, and emergency shutoff.[8] Stating requirements in that vocabulary lets security, facilities, and the AI team argue about the same objects.
What it does not change: where a control objective must be evidenced, the audit path stays the buyer's to own. Buying hardware before scoping the audit inverts the order of work.
HONEST LIMITS
When this is not the right fit
Modular infrastructure relieves a specific constraint. Where that constraint is not the binding one, it adds cost and complexity for nothing.
- Spiky or exploratory demand. If the cluster would idle for days at a time, elastic capacity is cheaper and faster. Stay there until the load flattens.
- Sub-megawatt requirements. A few racks of inference hardware belong in a colocation cage or an existing room, not in a megawatt-class unit.
- An existing facility with real headroom. If the hall still takes the power, the cooling, and the floor loading, expanding in place is the lower-risk path.
- No credible power path. A modular unit compresses construction, not interconnection; where energization is years out, the utility queue sets the schedule regardless.
- Compliance programs with no scoped audit path. Buying hardware before scoping the audit inverts the order of work.
- Teams that must always run the newest silicon. Owned capacity locks a generation for its depreciation life; rented capacity does not.
Before hardware selection
Ten checks, in clearing order
An unresolved item delays hardware selection rather than being worked around.
| # | Check | If it fails |
|---|---|---|
| 01 | Sustained utilization clears reserved cloud pricing over a quarter, not a peak week. | Stay on reserved cloud capacity. |
| 02 | A named residency, custody, or egress requirement rented capacity cannot satisfy. | The driver is preference — reassess. |
| 03 | A controlled site with a power path and a credible energization date. | Solve power before selecting hardware. |
| 04 | A rack-density target agreed by the AI team and facilities across the refresh horizon. | The envelope is wrong within a generation. |
| 05 | A cooling method matched to that density, with the heat-rejection path named. | Liquid-ready hardware lands in an air-only room. |
| 06 | Network design settled: one fabric or two, plus the wide-area path to the data. | The cluster waits on data instead of training. |
| 07 | Delivery route surveyed and permitting posture confirmed with the AHJ. | Schedule risk moves to the roadside. |
| 08 | An operating model on paper: monitoring, dispatch, spares. | Availability becomes an unowned problem. |
| 09 | Control objectives written as controls, with the audit path owned internally. | Nobody can define 'secure enough' at handover. |
| 10 | An exit and refresh plan for end of life or a workload that moves. | A five-year decision on an eighteen-month roadmap. |
In the product
How PODOS approaches enterprise deployments
PODOS builds capacity as a repeatable unit rather than a bespoke facility. Each PODOS Pod is designed as a standardized 1 MW building block and designed for 128 GPUs, with power distribution, closed-loop liquid cooling, and network interfaces integrated and tested before the unit leaves the factory. Capacity planning becomes arithmetic — units, not custom halls — and integration risk moves off the site. PODOS targets a 90-day window from order to commissioning for a standard unit; the sequence behind that target is on the deployment page.
PODOS AI is an early-stage company. Nothing here describes a completed deployment, a customer, or a certified product; the profile above is design intent. To size a configuration against a specific site, the configurator walks the same variables.
QUESTIONS
Frequently asked questions
When is dedicated AI infrastructure cheaper than cloud GPUs?
When utilization is high and sustained. Owned capacity is paid for whether or not it runs, so the comparison is the fully loaded cost of an owned megawatt divided by the hours it is actually busy, set against reserved cloud capacity. Workloads that idle most of the week rarely clear that bar; steady fine-tuning and production inference often do.
Does putting AI hardware on our own site make us compliant?
No. Physical custody changes where data sits and who can touch the hardware, which supports many control objectives — but compliance status comes from an audit against a named framework, not from an enclosure. PODOS AI claims no certification, attestation, or accreditation for any product.
What has to be true about the site first?
A power path with a credible energization date, a prepared pad with legal road access for the delivery, an authority having jurisdiction that accepts off-site-built equipment, and a heat-rejection strategy that fits the local climate and water position. Any one missing turns a hardware decision into a construction program.
Do we need a separate network for the GPUs?
For multi-node training, generally yes. Large-scale practice separates the frontend network from a dedicated non-blocking backend fabric carrying collective traffic between GPUs. Single-node or inference-only deployments can often stay on one fabric.
Sources
- [1] 2024 United States Data Center Energy Usage Report (LBNL-2001637) — Lawrence Berkeley National Laboratory, Dec 2024
- [2] Energy and AI — Executive Summary — International Energy Agency, Apr 2025
- [3] Global Data Center Survey 2025 — Uptime Institute, Jul 2025
- [4] GB200 NVL72 product page — NVIDIA, accessed 2026-08-31
- [5] Thermal Guidelines for Data Processing Environments, 5th ed. (TC 9.9) — ASHRAE, 2021
- [6] RoCE networks for distributed AI training at scale — Meta Engineering, Aug 2024
- [7] IEEE Std 802.3df-2024 — Ethernet Amendment 9 (800 Gb/s MAC; 400/800 Gb/s PHYs) — IEEE SA, published Mar 2024
- [8] SP 800-53 Rev. 5 — Security and Privacy Controls (PE control family) — NIST, Sep 2020, upd. 2025
- [9] Commercial Vehicle Size and Weight Program — federal size and weight standards — Federal Highway Administration (US DOT), accessed 2026-08-31
- [10] ICC/MBI 1205-2021 — Inspection and Regulatory Compliance in Off-Site Construction — International Code Council / Modular Building Institute, 2021 ed.
Size it against your site
Bring the utilization curve, the residency requirement, and the power path. The configurator walks the same variables an engineering review would.

