The Megawatt That Never Reached the Racks
The Megawatt That Never Reached the Racks

TL;DR

  • Plant validation does not equal rack validation: Using a temporary load to walk a coolant distribution unit (CDU) up to 1 MW confirms the primary facility loop and heat rejection, but it leaves downstream manifolds, hoses, quick disconnects, and cold plates completely unverified if the test loop bypasses them.
  • Trace the heat and scope the claim: Phased commissioning works only when the acceptance record explicitly follows the heat path from source to rejection, limiting pass claims strictly to the equipment the heat actually touched.
  • Set specific triggers for deferred testing: Vague terms like “verify at full load” create operational ambiguity; teams must establish concrete milestone triggers (such as 80% occupancy or a completed row) with defined entering conditions, datasets, and acceptance criteria while the team is still assembled.
  • Re-verify hydraulic and thermal balance as compute arrives: Each tranche of servers shifts the system curve, meaning operators must compare live production workloads against both initial baselines to confirm branch flow, valve authority, and thermal margins before signing off on full turnover

# # #

A temporary load can prove a liquid-cooling plant. It cannot prove equipment the heat never touched. Here is how to commission in phases without blurring that line.

By Gaurav Dhir, Chief Technology Officer, Reliability Engine

Imagine a liquid-cooled AI hall on the morning of turnover. The pumps are running, the coolant distribution unit (CDU) is stable, and the control graphics are green. The drawings show twenty 50 kW racks. Only four are on the floor.

To protect the schedule, the commissioning team connects a temporary thermal load beside the CDU and walks the system up to 1 MW. Temperatures settle. Pump speed holds. The trend plots look clean. The closeout report gets the sentence everyone has been waiting for: “Passed at 1 MW.”

That sentence may be accurate. It may also be misleading.

If the temporary circuit returned before the rack branches, the test heat never encountered the supply and return manifolds, flexible hoses, quick disconnects, or cold plates. The plant carried a megawatt. The rack path did not. Once the pass box is checked, that distinction is easy to lose.

There is nothing wrong with commissioning before every GPU arrives. The problem begins when the missing hardware also disappears from the acceptance record. The cleanest way to prevent that is simple: follow the heat from where it was created to where it was rejected, then limit the claim to the equipment it actually touched.

What Did the Megawatt Touch?

The answer depends on the type of load and where it was connected. An electrical load bank can exercise generators, UPS systems, switchgear, power distribution, and room-level heat rejection. It does not place heat into a direct-to-chip loop.

A thermal load connected to the technology cooling system (TCS) near the CDU can test heat exchanger duty, pump operation, control stability, and the response of the facility-water loop. If it returns through a short temporary circuit, it says little about pressure loss or flow distribution in the rack branches.

Move that load to a rack manifold and the boundary becomes more useful. The test can now include permanent branch piping, valves, hoses, and local instruments. It still cannot reproduce cold plate pressure drop, the liquid-to-air heat split, server controls, or the timing of a real AI workload.

The temporary hardware matters as well. A portable heat exchanger, clean filter, or oversized hose can create an easier hydraulic path than the permanent installation. The report should show the connection points, identify temporary components, and mark the permanent equipment that remained outside the test.

The megawatt is the headline. The flow path is the story. Without both, the result is too easy to misread six months later when the people who ran the test are no longer in the room.

Prove the Plant, Then the Path

Before the first production rack carries load, finish the work that establishes a trustworthy fluid boundary. Confirm pressure integrity, valve lineup, cleaning and flushing, air removal, filtration, coolant chemistry, sensor mapping, leak detection, alarms, and control sequences. A good load test cannot rescue a loop that was never properly prepared.

Next, test the plant at the low, partial, and design conditions defined in the commissioning plan. Record temperatures on both sides of the CDU, TCS flow and differential pressure, pump command and feedback, valve or bypass position, facility-water conditions, and the control mode in force. Include redundancy and failure scenarios, but only claim branch recovery when the test connection exposed the limiting branch.

When the first production racks arrive, the test changes character. Use a repeatable workload to establish the operating baseline, then follow its heat all the way to final rejection. Record rack power, branch flow, differential pressure, coolant temperatures, CDU behavior, relevant air-side conditions, and the server telemetry used for acceptance.

The measurements will not agree perfectly. Some heat still goes to air, instruments have uncertainty, and the system stores heat during a ramp. They should still tell a coherent story. If rack power rises while liquid-side heat stalls, or pump demand climbs at the same thermal duty, the gap needs an explanation before the result is accepted as normal.

Choose the first racks carefully. The rack nearest the CDU is usually the easiest hydraulic case, not the most informative one. Include a long branch, a restrictive path, a different elevation, or another hardware configuration when those differences exist.

The work does not end with the first four racks. Each new tranche changes the system curve and can shift flow toward easier paths. There is no need to repeat the entire program after every delivery. Define sensible checkpoints, then verify branch flow, valve authority, pump operation, alarm mapping, and heat balance as the hall fills.

Write What Actually Passed

Return to the illustrative twenty-rack block. The temporary load demonstrates that the CDU, facility-water loop, and heat rejection plant can carry 1 MW at the documented entering conditions. The four live racks demonstrate the production path only at their locations and workloads. Final branch balance remains open.

That is not a failed commissioning program. It is a phased one. The difference lies in how the result is written.

An acceptance statement I would be comfortable signing might read: “CDU Block A sustained 1 MW of thermal duty at the documented entering conditions. Four identified production rack paths were verified at their recorded workloads. Final branch-balance verification is deferred until the block reaches the agreed occupancy trigger.”

That sentence is longer than “1 MW passed,” but it is far more useful. The operator can see what the plant proved, what the production racks proved, and which question still has no answer. The next team does not have to reverse-engineer the test from trend files and temporary piping photographs.

Make Later Specific

Deferred testing often fails in the wording. “Verify at full load” sounds reasonable until nobody can agree on what full load means, who owns the test, or which data will close it.

Name the trigger while the project team is still together. It might be 80 percent block population, the first complete row, the arrival of a new rack family, or the first planned redundant-mode run. Define the load source, duration, entering conditions, control mode, redundancy state, data set, acceptance criteria, approver, and response to a failed result.

Prepare the data path at turnover, not on the morning of the deferred test. Check timestamps, units, sensor identity, trend intervals, and command and feedback points. An unsynchronized rack-power log or compressed historian trace can make a sound physical test impossible to interpret later.

Keep the temporary-load baseline and the production-IT baseline separate. Both are valuable, but they describe different hardware and flow paths. Combining them into one normal range erases the very difference the phased program is meant to preserve.

When the Compute Finally Arrives

Before running the deferred test, confirm that the physical installation still matches the turnover record. Check valve lineup, connectors, sensor mapping, filter condition, makeup-fluid history, and any construction or service work completed during the gap. A hall can change considerably while everyone waits for servers.

Run the documented workload and compare it with both earlier baselines. The traces will differ because occupancy, workload, hardware, heat split, and controls have changed. Those differences should be explainable. If they are not, look for a restriction, imbalance, measurement problem, control issue, or loss of thermal margin.

Only then close the deferred item. The temporary load showed what the infrastructure could carry. The early racks showed how selected production paths behaved. The populated hall shows whether those paths still work together under the configuration the owner is about to operate.

The Number Is Not the Proof

A liquid-cooled AI hall can open before every GPU arrives. What it cannot afford is an acceptance record that lets absent hardware masquerade as tested hardware.

The most useful question is not only, “How much load did we test?” It is, “Where did the heat go?” Follow that path, record what it touched, and keep the remaining test visible. Then the megawatt in the report means exactly what the operator thinks it means.

# # #

About the Author

Gaurav Dhir is Chief Technology Officer at Reliability Engine, where he works on predictive reliability for liquid-cooled AI infrastructure. His focus includes coolant condition, hydraulic behavior, commissioning evidence, and the operating data needed to protect GPU capacity. He also writes Reliability Engine Insights on practical liquid-cooling reliability.

References

  1. ASHRAE. AI Data Center Energy Performance Framework: Commissioning and Performance Validation. 2026.
  2. ASHRAE Handbook, HVAC Applications, Chapter 20: Data Centers and Telecommunication Facilities. Section on commissioning and load bank planning.

The post The Megawatt That Never Reached the Racks appeared first on Data Center POST.


Discover more from Website Hosting Review

Subscribe to get the latest posts sent to your email.