Andris Gailitis

Cloud, Colo & Data Center Pro

Tier 6? When the Server Demands More Redundancy Than the Building

Tier 6? When the server demands more redundancy than the building - AI data center with liquid cooling pipework and multi-layer redundancy
  • Key takeaways: The server is now designing the building. AI racks ship at 120-135 kW today, up to 180 kW on the newest platforms – and the facility has to follow, not the other way round.
  • Six power feeds on a rack do not mean six UPS paths. An NVL72 rack carries 48 PSUs across 8 power shelves, aggregated by bus bar into far fewer upstream paths.
  • Above roughly 400 kW per rack, 415 V AC stops making sense. NVIDIA has announced 800 VDC architecture for AI factories from 2027.
  • What becomes obsolete: centralised UPS, static switchgear, fixed distribution. What becomes critical: modular UPS, BESS, distributed power, CDUs, high-bandwidth fabric.
  • Evidence base: 475 signals from 87 sources, plus a full 20 MW reference architecture for 2028 delivery – free PDF attached.

For thirty years we built data centers to house servers. The building came first, the racks moved in afterwards, and the tier level on the certificate told you how good the building was. That order has now reversed. The server is designing the building.

We spent weeks pulling apart what current AI platforms actually demand from a facility – 475 signals from 87 sources – and the picture that comes out is not an incremental step. It is a different machine.

From 10 kW to 180 kW in one hardware cycle

A conventional enterprise rack sat at 5-10 kW for most of my career. Here is where the shipping AI platforms are now:

PlatformRack powerPower feedsCoolingLiquid share
NVIDIA GB200 NVL72120 kW6Direct liquid100%
NVIDIA GB300 NVL72135 kW6 or 8Hybrid90%
AMD Helios (MI455X)up to 180 kWDirect liquid100%

The forecast in our data set puts typical racks at 300-400 kW by 2030, with high-end configurations reaching 480 kW. If you are commissioning a building today for a 2028 delivery, that is the load you are designing for – not the one in front of you.

Does six power feeds mean six UPS systems?

This is the question I get asked most often, and the honest answer is no – but you have to look inside the rack to see why. An NVL72 rack carries eight power shelves, six PSUs each at 5.5 kW: 48 power supplies in a single rack, with N+N redundancy at shelf level. Those shelves feed a bus bar, and the bus bar feeds a remote power panel through power whips. The redundancy is real, but it lives inside the rack and collapses into far fewer paths upstream.

Which means the old habit of counting cords to judge resilience no longer tells you anything useful. A rack with six inlets can still be sitting on one failure domain.

Where the electrical architecture is going

415Y/240V distribution is becoming the default for new AI facilities, replacing the older 480/277V standard. But that is a transition step, not a destination. Above roughly 400 kW per rack the conversion losses stop being acceptable, and the evidence points one way: NVIDIA has announced a move to 800 VDC architecture for AI factories starting in 2027, with integrated multi-time-scale energy storage.

The safety objection people raise first – arc flash – has an answer in the data: properly designed 800 VDC systems reach arc flash safety levels comparable to traditional AC data center power.

Alongside that, the UPS question changes shape. AI loads do not draw smoothly. Training jobs create rapid, correlated power swings that push UPS systems into frequent micro-cycling and stress the distribution behind them. That is why battery energy storage moves from nice-to-have to structural – typically sized for 1-4 hours, not for the seconds you need to start a generator. And the most interesting idea in the whole data set: the training scheduler itself could eventually become part of the power-control system, shaping demand before it reaches the switchgear.

How much air is left?

Very little. By 2028 residual air cooling is expected to handle only 5-10% of the heat, with liquid taking the rest. The design centre of gravity moves to the CDU: Schneider’s MCDU-70 delivers 2.5 MW per unit at an industry target flow of 1.5 litres per minute per kW; CoolIT’s CHx2000 supports up to 2,000 kW. For a 10 MW facility, six MCDU-70 units in a 4+2 configuration is a documented answer.

Note what that does to your failure domains. Each CDU serves a row or group of racks. Your cooling topology is now a redundancy topology, and it does not necessarily line up with your electrical one.

Tier 6? What the tier language stops describing

Traditional tier classification describes building-level redundancy: paths, distribution, concurrent maintainability. It was built for a world where the rack was a passive container.

An AI rack is not passive. It has its own power redundancy, its own cooling loop, its own failure domain, its own control logic. Redundancy in these facilities lives across at least ten layers – power distribution, cooling, network fabric, server hardware, facility plant, energy storage, control systems, monitoring, security, maintenance discipline – and a building-level tier badge says almost nothing about most of them.

I am not proposing a literal Tier 6, and neither is the report. The point is sharper than that: we are still describing these facilities with vocabulary that was designed for a different machine. A building can hold its certification and still be the wrong shape for the load inside it.

What dies, what matters

Becoming obsoleteBecoming critical
Centralised UPS systemsModular UPS
Static switchgearBESS / energy storage
Fixed distribution infrastructureDistributed power architecture
Air as the primary cooling pathCDUs and liquid loops
PUE as the headline metricUseful compute per MW

That last row deserves its own article. PUE measures how much power you waste on everything that is not compute. It says nothing about whether the compute itself is doing useful work. In an AI facility, two buildings with identical PUE can differ enormously in delivered training throughput per megawatt – and that is the number the customer is actually buying.

The 30-year building and the 2-year server

This is the operator’s real problem. We finance and build shells that must last three decades, to house IT that turns over every two years, at densities we cannot yet confirm. The answer in the data is not heroic over-provisioning – that is just waste capex with a nicer name. It is modularity: standardised power building blocks from 500 kW upwards, with defined electrical, mechanical and control interfaces, and layouts that can absorb a density step without a rebuild.

The report closes with a full 20 MW reference architecture for 2028 delivery – 167 racks at 120 kW, 415Y/240V distribution, N+1 UPS per row, CDUs in 4+2, mapped failure domains – and what changes when you scale it five times to 100 MW: medium-voltage distribution, a grid connection up to 36 kV, water management, heat rejection. It is a design starting point, not a product recommendation, and every assumption is written down.

📄 Full research report (475 signals, 87 sources, 12 datasets, 20 MW and 100 MW reference architectures – free PDF, no paywall): AI Data Center Architecture Research 2026

If you are designing or buying capacity for 2028: what density are you actually specifying – and what happens to that building if the number doubles? 👇

#AIInfrastructure #DataCenters #LiquidCooling #PowerInfrastructure #GPUClusters

Also published in my LinkedIn newsletter: Cloud, Colo & Data Center Pro

Related articles: