By Chander Kant Updated July 31, 2026
Chapter 2 followed power and cooling to the boundary of one node. We now follow those connections through the rack and into the building. A power supply cannot draw current that the branch circuit cannot deliver, and a fan cannot remove heat into air that is already too warm. The facility is part of the computing system because every calculation ends by placing heat into it.
This relationship is easiest to see when a cluster is installed in an existing machine room. The server quote may fit the budget and the rack count may fit the available floor area, yet the project can still fail because the room cannot deliver enough power to those rack positions, cannot remove the resulting heat, or cannot carry the equipment safely through the loading dock and onto the floor. Dense accelerator racks make these constraints more visible, but they apply to ordinary CPU clusters as well.
We will begin with measured electrical load, follow power through the distribution system, and then trace the heat through air or liquid cooling. The final sections cover racks, floor loading, cabling, service space, instrumentation, and staged commissioning. The aim is not to turn the cluster designer into an electrical or mechanical engineer. It is to provide enough understanding to state requirements, recognize coupled decisions, and test the completed installation with the facility professionals who own those systems.
Establishing the design load
A facility design begins with a load estimate. The useful estimate is not the sum of optimistic idle measurements, nor is it automatically the sum of every nameplate rating. Nameplates protect equipment and distribution paths against specified limits. Application measurements describe what a particular configuration actually draws while running. The two numbers answer different questions, and a responsible design records both.
Measure representative nodes at idle, during boot, and under CPU, memory, accelerator, network, and storage stress. Measure switches and storage systems as well. A workload may not reach every component's individual maximum at the same instant, but synchronized applications can create broad load changes across many nodes. Distributed training, for example, can move repeatedly between communication-heavy and compute-heavy phases. Power caps and clock controls can bound some of this behavior, but the cap must be tested as a system property rather than assumed from a configuration screen.
The estimate also needs a boundary. Decide whether a rack figure includes only compute nodes or also top-of-rack switches, cooling pumps, coolant distribution units, storage, management equipment, and conversion losses. A 50 kW compute load and a 50 kW facility feed are not equivalent if the feed must also supply several kilowatts of supporting equipment.
| Evidence | What it establishes | Common mistake |
|---|---|---|
| Equipment nameplate and vendor site guide | Supported input, connectors, maximum current, cooling conditions, and installation limits | Treating the maximum rating as a prediction of normal workload power |
| Node measurements under representative work | Idle, typical, peak, and phase-dependent demand for the selected configuration | Testing one convenient kernel that does not exercise memory, accelerators, or I/O |
| Simultaneous rack test | Aggregate demand, load steps, phase balance, power-cap behavior, and thermal interaction | Multiplying one-node averages without checking synchronized behavior |
| Growth and failure allowance | Capacity needed for expansion, maintenance states, and surviving a failed redundant component | Using every available ampere on installation day |
| Facility measurement boundary | Which conversion, pump, fan, and supporting loads are included | Comparing numbers taken at different points in the electrical path |
Suppose sixteen nodes draw 2.6 kW each during the selected acceptance workload, two switches draw 1.2 kW each, and management and local storage draw another 3 kW. The measured IT load is 47 kW. Adding 15 percent for approved growth gives a planning figure of about 54 kW. This is an illustrative calculation, not a breaker size. The electrical engineer still has to apply the equipment ratings, continuous-load rules, redundancy model, ambient conditions, conductor limits, protective devices, and local code.
Current rack-scale systems show why measurement boundaries and product generations matter. NVIDIA lists 120 kW as the designed rack power for GB200 NVL72 and up to 142 kW for GB300 NVL72. The systems include compute trays, NVLink switch trays, management switches, power shelves, a busbar, and cooling manifolds, but their requirements are not identical. These figures are useful examples of present rack density. An installation still has to use the site guide, bill of materials, workload measurements, and redundancy state for the exact equipment being purchased.
Electrical distribution
Electrical power may pass through utility service, switchgear, transformers, uninterruptible power supplies, distribution panels, power distribution units, rack busbars or outlet strips, server power supplies, and voltage regulators before it reaches a processor. Every stage has a voltage and current rating, a conversion efficiency, a protective device, and a failure boundary. Figure 5.1 shows a simplified path with two independent rack feeds.
There is more than one way to build the last part of that path. A conventional rack may distribute AC through rack PDUs to a power supply in each server. Open Rack designs and integrated high-density systems may instead use centralized power shelves to convert AC to roughly 48 to 51 V DC, then carry the DC through a rack busbar. NVIDIA's DGX Grace Blackwell rack is one current example. This arrangement moves the conversion and redundancy boundaries into the rack design, so the power shelves, busbar, controls, and service procedure have to be evaluated with the compute trays.
Voltage, current, and three-phase service
For a single-phase AC load, real power is approximately voltage multiplied by current and power factor:
P = V I PF
For a balanced three-phase load expressed with line-to-line voltage, the corresponding approximation is:
P = sqrt(3) VLL I PF
Power factor describes how much of the apparent power, measured in volt-amperes, becomes real power, measured in watts. Modern server power supplies usually correct their power factor over the supported operating range, but the project should use the vendor's data rather than assume a value. Conversion efficiency also changes with load. A redundant power supply carrying a small fraction of its rating may operate at a different efficiency from the same unit near its preferred range.
The formulas are useful for checking an estimate. On a 415/240 V wye service, a balanced 54 kW load at 415 V line-to-line and power factor 0.95 draws about 79 A per phase. At 208 V, the same load and power factor would draw about 158 A per phase. Use either result only as an estimate to reconcile with branch ratings, continuous operation, protective coordination, conductor temperature, connector limits, and the effect of losing one power path.
Three-phase distribution also creates a balancing problem. A rack PDU may present three groups of outlets tied to different phases. Placing all high-load nodes on one group can overload a phase while the aggregate rack power appears to be below its limit. Meter each phase, document which outlets belong to it, and preserve that mapping when equipment is moved.
Redundancy has to survive the intended fault
Terms such as N+1, 2N, A/B, and concurrently maintainable describe different arrangements. They are not interchangeable assurances that the cluster will stay up. N+1 means that one additional unit is available beyond the number required for the design load. A 2N system duplicates the required capacity. A/B rack feeds normally refer to two distribution paths. The actual result depends on where those paths begin, which components they share, and how loads transfer when one side fails.
A node with two power supplies should normally connect one supply to each independent feed. Both feeds must be able to support the agreed failure state. If each supply shares half the load during normal operation, losing one feed can nearly double the demand on the other. A design that works only while both sides are healthy is dual-fed, but it is not redundant at the stated load.
Trace common points all the way upstream. Two rack PDUs connected to the same panel, UPS module, transfer switch, transformer, or generator may protect against a failed outlet strip but not against the larger event the team had in mind. Cooling and control systems need the same examination. Redundant compute feeds do little good if a single failed pump, network switch, or control power supply removes cooling from the rack.
UPS capacity covers a defined interval and set of disturbances. A generator covers a different interval, after starting and accepting load. Neither should be described simply as "backup power." Record ride-through time at the design load, battery condition, generator start and transfer sequence, fuel assumptions, maintenance bypass behavior, and which cooling components remain powered. Then test the sequence without risking an uncontrolled shutdown of the full cluster.
Power sequencing and load steps
A large cluster should not necessarily start every node at once. Simultaneous power-supply inrush, fan acceleration, storage spin-up where rotating media remains, network boot, and application restart can create a larger step than normal computation. Staged startup reduces the step and makes failures easier to isolate. It also prevents thousands of nodes from asking DHCP, provisioning, authentication, and monitoring services for work at the same instant.
Shutdown needs an order as well. Stop or checkpoint applications, drain scheduler allocations, protect unique data, quiesce storage, and confirm that dependent control services will remain available. An emergency power-off system has a different purpose: protecting people and the facility under conditions where an orderly application shutdown may be impossible.
Power caps can enforce a facility boundary, but a cap changes application behavior. A node that reaches its cap may lower CPU or GPU clocks, producing longer runtimes and changing when communication and I/O occur. Measure job completion, not only instantaneous watts. A lower cap may improve jobs per unit of energy when performance drops slowly; it may reduce useful throughput when a synchronized job waits for capped ranks.
Power quality and disturbance measurement
Average power does not describe every electrical interaction. Server power supplies draw switched current, accelerator workloads can change demand quickly, and thousands of devices can make similar transitions at nearly the same time. The distribution system has to keep voltage within the equipment envelope while protective devices distinguish a real fault from an expected load step.
Measure voltage events, current by phase, frequency, power factor, and harmonic behavior at the boundary appropriate to the problem. A disturbance visible at one rack but not an adjacent rack points to a different path from an event seen across the room. Time-aligned node and scheduler records help identify whether the event coincided with boot, application launch, a collective phase, a power-cap transition, or a facility transfer.
Protective coordination determines which breaker or fuse opens for a fault. The desired behavior is usually to isolate the smallest affected portion without allowing upstream equipment to remain exposed. That analysis belongs to the electrical engineer and the approved equipment design. The cluster team's contribution is an accurate load profile, startup sequence, rack map, and list of computing states that must be tested.
Every watt becomes heat
Nearly all electrical power entering IT equipment becomes heat in the room or cooling loop. A 54 kW rack therefore presents approximately a 54 kW thermal load while it is drawing that power. In customary HVAC units, one watt is about 3.412 British thermal units per hour, so 54 kW is roughly 184,000 BTU/h. The conversion does not include a safety margin; it only expresses the same heat flow in another unit.
Cooling must remove heat at the rate it is produced. For a fluid, the transferred heat can be written as:
Q = mass flow x specific heat x temperature rise
When flow is measured by volume, mass flow equals density multiplied by volumetric flow. This relationship helps explain why a practical air-cooled rack needs a large volume of air and a controlled path through the equipment, while a liquid loop can carry the same heat in a much smaller volume. Liquid cooling adds pumps, heat exchangers, hoses, seals, water chemistry, and another set of failure modes. Air and liquid cooling therefore have to be compared as complete systems.
Airflow and containment
Air-cooled servers usually draw cool air through the front and discharge warmer air at the rear. The temperature that matters most to the server is the inlet temperature at its air intake. A thermostat several meters away can report an acceptable room average while the top of a rack recirculates hot exhaust into its own inlet.
Hot-aisle and cold-aisle arrangement separates supply air from exhaust. Containment strengthens that separation. Blanking panels close unused rack positions so air cannot bypass equipment. Sealing cable openings and managing pressure reduces short circuits in which cold air returns to a cooling unit without passing through a server, or hot air returns to an inlet without reaching a cooling unit.
More fan speed cannot repair every airflow problem. Fans operate against resistance. Dense cables, closed doors, filters, heat sinks, and server chassis all contribute pressure drop. If the room cannot supply the required air at the required pressure, server fans consume more power and still see a high inlet temperature. Measure temperatures at several rack heights, differential pressure where relevant, fan state, and equipment throttling. Correlate these measurements with IT load rather than inspecting cooling on an idle Sunday.
Temperature and humidity limits come from the installed equipment and the facility design. ASHRAE's data-center guidance distinguishes recommended conditions used for efficient, reliable operation from wider allowable envelopes in which qualified equipment can function. Running inside an allowable edge is not the same as choosing it as the normal operating point. Humidity, dew point, electrostatic risk, particulate contamination, altitude, and equipment class all affect the decision.
Altitude deserves an explicit check. Air density falls as elevation increases, so the same volumetric airflow carries less mass and removes less heat. Servers, fans, air handlers, and other cooling equipment may have altitude limits or temperature derating rules. Use the limits for the selected equipment and the actual site instead of assuming that a sea-level cooling test will transfer unchanged.
Air cooling reaches a practical limit when the required volume, pressure, fan energy, heat-exchanger area, and aisle geometry no longer fit the intended rack density. That point differs by building and equipment. Rear-door heat exchangers can extend an air-cooled arrangement by removing heat from rack exhaust. Direct liquid cooling moves heat from selected components before it enters the air. Many high-density systems are hybrid: processors, accelerators, and switches may be liquid cooled while memory, power supplies, storage, and other components still heat the air.
| Architecture | Principal heat path | Questions that decide the fit |
|---|---|---|
| Room air with contained aisles | Server fans move room supply air through the chassis to the return-air system | Required airflow and pressure, inlet uniformity, fan energy, rack density, and room cooling capacity |
| Rear-door heat exchanger | Server exhaust passes through a liquid-fed door before returning to the room | Door weight and clearance, water connections, fan interaction, condensation control, and heat captured at peak load |
| Direct-to-chip liquid cooling | Cold plates move component heat into a TCS loop; air removes the remaining load | Liquid capture ratio, coolant temperature and chemistry, flow, pressure, CDU redundancy, hoses, and leak response |
| Immersion | Equipment transfers heat to dielectric fluid and then to a heat exchanger | Compatible hardware and materials, tank operation, fluid handling, maintenance, cabling, fire review, and warranty |
Direct liquid cooling
Direct-to-chip cooling places cold plates on high-heat components and circulates coolant through them. A rack manifold distributes coolant to equipment. The technology cooling system, or TCS, is the loop serving the IT equipment. A coolant distribution unit, or CDU, commonly uses a heat exchanger to transfer heat from the TCS to the facility water system, or FWS, while keeping the fluids separate. The facility loop then rejects or reuses the heat.
The Open Compute Project's liquid-to-liquid CDU test methodology describes the CDU as the equipment separating the TCS from the FWS. That boundary protects the IT loop from facility pressure, chemistry, and contamination, and lets the CDU regulate flow and temperature for the installed cold plates. It also means that both sides of the heat exchanger have finite approach temperature, flow, and pressure-drop requirements.
Liquid-to-air CDUs provide another arrangement for sites without facility water. They circulate the TCS coolant but reject its heat into the room air, so the room cooling system must carry that load. The distinction belongs in the design documents: a liquid-cooled server connected to a liquid-to-air CDU has changed the heat path inside the rack, but it has not removed the heat from the air-cooled facility.
Warmer supply coolant can reduce or eliminate compressor work in a suitable climate and can make recovered heat more useful. The equipment must be qualified for the selected coolant temperature, and the supply must remain above the local dew point where exposed surfaces could condense moisture. A warm-water label does not by itself prove that chillers are unnecessary; outside conditions, required approach temperatures, redundancy, and the facility heat-rejection design decide that.
Flow, chemistry, and failure behavior
A liquid-cooled rack arrives with requirements for flow, supply temperature, pressure, pressure differential, coolant composition, filtration, materials, and allowable transients. These values are part of the equipment interface. Increasing pump speed to cure a hot component can exceed a pressure limit or move a pump out of its intended range. Reducing flow to save pump power can create a local thermal limit before the rack-level return temperature looks unusual.
Material compatibility and water chemistry affect long-term reliability. Mixed metals, unsuitable elastomers, oxygen, particles, biological growth, and incorrect additives can corrode or obstruct small passages. The facility should have a documented fill, flush, filter, sample, and maintenance process. Commissioning debris that would be unimportant in a large pipe can damage a cold plate or quick disconnect.
Quick disconnects and hoses need support, correct bend radius, inspection access, and a defined replacement practice. Dripless does not mean incapable of leaking. Place leak detection where fluid would actually travel, connect it to an alarm and isolation response, and test the response. Decide what happens after loss of a pump, CDU controller, facility-water flow, or one electrical feed. The node firmware may throttle or shut down, but the time available depends on thermal mass and residual heat.
The air side remains part of the calculation. If cold plates remove 80 percent of a 100 kW rack load, the room still has to remove 20 kW plus any CDU or power-conversion heat discharged into the space. Vendor documentation should state the liquid heat-capture ratio under the intended configuration. Verify it at load, because missing cold plates, different components, and warmer local air can change the remainder.
Immersion cooling places equipment in a dielectric fluid and can remove a large fraction of the heat without conventional server airflow. It changes server form factor, maintenance procedures, material compatibility, cabling, fire review, fluid handling, warranty, and the way failed components are removed. Fluid supply, disposal, and changing environmental requirements also belong in the lifecycle review. These changes make immersion a platform decision for the facility and its equipment.
Energy and water metrics
Power usage effectiveness, or PUE, is the ratio of total annual facility energy to annual IT equipment energy:
PUE = total facility energy / IT equipment energy
A perfect ratio of 1.0 would mean that every unit of facility energy reached IT equipment and none was used for cooling, conversion, lighting, or other support. Real facilities are above 1.0. PUE is useful for tracking one facility with a stable measurement boundary. It does not report how much useful computation the IT equipment completed. A lightly utilized cluster can sit in an efficient building and still waste more energy per completed job than a busy cluster in a building with a higher PUE.
Water usage effectiveness, or WUE, commonly reports annual site water use divided by annual IT energy, often in liters per kilowatt-hour. The result depends on what water and which site boundary are counted. A dry cooler can reduce on-site water consumption while the electricity it uses has water and carbon consequences elsewhere. A cooling tower withdraws water into the site and consumes the portion that evaporates or leaves the local watershed. The balance may also include blowdown returned to treatment. Withdrawal and consumption are therefore not synonyms.
Record the measurement point, interval, climate, load, and included systems beside any PUE or WUE value. Do the same for carbon and heat-reuse metrics. Reusing warm return water can be valuable when a nearby load needs heat at the available temperature and time. A theoretical heat quantity is not recovered energy unless a practical consumer, heat exchanger, pumping path, and operating schedule exist.
Facility efficiency should be connected to computing output. Jobs completed per kilowatt-hour, simulations per unit of energy, training progress per unit of energy, or another workload-specific result exposes changes hidden by PUE. Result quality belongs in the denominator or acceptance condition when lower precision, smaller models, or early termination change the answer.
Racks, floors, cables, and service space
A rack is both a mechanical structure and an interface to the room. Its width, depth, height, mounting standard, power distribution, cooling connections, door perforation, cable space, and maximum load must match the equipment. Traditional 19-inch racks, Open Rack designs, busbar systems, and integrated rack-scale products are not interchangeable merely because they occupy a rectangular floor area.
Floor loading and equipment movement
Check the complete installed weight: rack, servers, switches, power shelves, batteries if present, coolant, manifolds, doors, cables, and any shipping or service fixture. A structural engineer must evaluate the building. Raised floors add tile, pedestal, stringer, rolling-load, concentrated-load, and underfloor-obstruction questions. Slab floors still have concentrated and rolling limits.
The route to the rack can be more restrictive than the final position. Check loading-dock capacity, door width and height, corridor turns, ramps, thresholds, elevator dimensions and rating, floor transitions, and overhead clearance. A floor that supports a stationary rack may need spreader plates to carry a heavy rack on casters. Integrated liquid-cooled racks can also require a filling or final assembly sequence that changes their transport weight.
Center of gravity matters during movement and installation. Heavy equipment should be placed according to the rack and vendor instructions, commonly keeping the rack stable and avoiding a top-heavy arrangement. Leveling feet, anchoring, seismic requirements, and joining adjacent racks depend on the site. The Open Compute Project's colocation guidance distinguishes rolling, uniform, and concentrated floor loads for this reason; one floor rating does not answer all three.
Rack layout and cabling
Rack elevations should show equipment, blanking panels, power distribution, switches, patch panels, manifolds, CDUs, cable managers, and reserved space. Place components with airflow and service in mind, not only to minimize cable length. A failed power supply, fan tray, switch, or hose must be removable without disconnecting unrelated equipment.
Cables have minimum bend radii and pull limits. High-speed copper and optical assemblies can be damaged without visible external evidence. Support their weight, avoid sharp turns near connectors, use hook-and-loop rather than crushing ties where appropriate, and keep bundles out of fan intakes and equipment-removal paths. NVIDIA's data-center cabling guidance recommends large loops for managed slack, individual path control, proper support, and never routing cables through the middle of a rack where they restrict airflow and service.
Label both ends with identifiers tied to an authoritative port map. Record cable type, length, source and destination, transceiver, rail, and installation test. Chapter 3 explained that topology affects performance; the physical cable record is how that logical topology can be verified and repaired. Separate network scopes visually or by managed pathways where practical, and protect management and storage cabling from accidental changes to the compute fabric.
Cooling hoses need the same discipline. They should not carry their own unsupported weight, rub against sharp edges, obstruct electronics, or require excessive force at a quick disconnect. Distinguish supply and return, label manifold positions, record the served equipment, and leave enough controlled movement for service without creating a trip or snag hazard.
Service space and safety
Service clearances are working space, not spare floor area. Technicians need room to open doors, slide equipment, operate lifts, attach test instruments, contain a leak, and remove a failed component. Staging space is needed for unpacking, inspection, burn-in preparation, and packaging of returns. Spare parts need controlled storage and inventory rather than an improvised pile in the machine room.
Dense clusters can produce hazardous sound levels. Heavy components create lifting and crush risks. Electrical distribution, batteries, cooling fluids, raised floors, fire detection and suppression, and emergency power-off systems all have site-specific safety requirements. Energized-equipment working clearances and arc-flash assessment come from the approved electrical design and local authority; an aisle that is wide enough for server service may not satisfy them. These matters belong with qualified facility and safety professionals. The cluster team still needs to understand them because maintenance procedures, staffing, remote operation, and equipment selection depend on the result.
Fire and leak responses should preserve life safety first and define what happens to computing equipment second. Decide which alarms drain jobs, which trigger an orderly shutdown, which isolate coolant, and which remove power immediately. Avoid automation that turns a faulty sensor into an unnecessary facility-wide outage, but do not require a human to investigate a dangerous condition before protective action can occur.
Facility and cluster telemetry
Power and cooling systems already contain meters and controllers. The cluster contains node, switch, storage, scheduler, and application measurements. Their clocks and identifiers should be aligned well enough to reconstruct an event. A job slowdown may coincide with a power cap, high inlet temperature, failed pump, network error, or a facility transfer. Separate dashboards with incompatible timestamps make the relationship harder to see.
Useful electrical measurements include real power, apparent power, current by phase, voltage, power factor, breaker state, UPS state, battery condition, generator events, and transfer events. Thermal measurements include rack inlet and exhaust temperatures, humidity or dew point where relevant, coolant supply and return temperatures, flow, pressure differential, pump state, fan state, leak alarms, and CDU or facility alarms.
Set alarms from equipment limits and operational experience. An alarm should identify a responsible team and a response. Warning thresholds can allow a job drain or capacity reduction; critical thresholds may require rapid shutdown. Review alarms that operators repeatedly ignore, because that background noise can hide the event that matters.
Retain enough history to compare seasons, workloads, rack populations, and component aging. A gradual rise in pressure drop may reveal filtration or flow trouble. Increasing fan power may show airflow obstruction. A growing difference between redundant feed currents may reveal failed power supplies or phase imbalance. These trends inform Chapter 8's maintenance and repair process and the capacity and expansion decisions discussed later in Chapter 10.
Commissioning and acceptance
Commissioning should establish evidence before expensive equipment depends on the result. The exact procedure belongs to the project design and qualified facility team, but a cluster acceptance sequence normally proceeds from static inspection to controlled load and then to failure behavior.
- Verify documents and labels. Reconcile rack elevations, one-line electrical diagrams, panel schedules, breaker and outlet labels, cable maps, coolant diagrams, valve positions, alarm routes, equipment ratings, and the final bill of materials.
- Inspect the physical route. Confirm access, floor protection, rack anchoring or leveling, clearances, airflow direction, blanking, cable support, hose support, bend radii, leak detection, fire systems, and safe service access.
- Test distribution without the cluster load. Exercise meters, protective devices, control power, UPS and generator interfaces, cooling controls, pump redundancy, valves, alarms, and communications according to the approved commissioning plan.
- Introduce load in steps. Power management equipment first, then a small set of nodes, a rack, and finally the intended cluster scope. Observe voltage, phase current, power factor, inrush, cooling response, inlet temperatures, coolant flow, pressure, and leaks at each stage.
- Run representative sustained work. Exercise CPU, memory, accelerators, network, and storage long enough for temperatures and cooling controls to settle. Confirm that no component throttles unexpectedly and that the measured design load agrees with the planning model.
- Create controlled transitions. Test startup sequencing, orderly shutdown, power caps, scheduler drain, and restoration. Confirm that provisioning and control services tolerate the selected startup batch size.
- Exercise approved failure cases. Remove one redundant feed, pump, cooling component, or control path only under the documented safe procedure. Confirm surviving capacity, alarm delivery, automatic action, operator response, and recovery.
- Record the baseline. Preserve configurations, meter boundaries, acceptance workloads, environmental conditions, measurements, alarms, exceptions, and sign-off. Future changes should be compared with this known state.
The test should include the application boundary. A facility can hold temperature while a power cap silently reduces throughput. A rack can remain below its aggregate current limit while one phase or connector is overloaded. A CDU can report adequate total flow while one branch is restricted. Acceptance succeeds only when the component measurements and the completed workload tell the same story.
Before the first job
A useful final review follows the complete paths rather than checking power and cooling as separate totals. Establish where power is measured, which components survive a failure, where heat enters each cooling loop, how the remaining air load is removed, and how people will reach the equipment when something fails. Then test those paths at representative load.
Chapter 6 moves from the physical machine to the Linux and application software environment. The facility work remains visible there: processor frequency, accelerator availability, network behavior, storage performance, and reproducibility all depend on a stable physical platform underneath them.
References and further reading
- U.S. Department of Energy, Best Practices Guide for Energy-Efficient Data Center Design, revised 2024.
- U.S. Department of Energy, "Cooling Water Efficiency Opportunities for Federal Data Centers."
- ASHRAE, "Energy and Thermal Efficiency," AI Data Center Energy Performance Framework.
- ASHRAE Handbook, "Data Centers and Telecommunication Facilities."
- Open Compute Project, Liquid-to-Liquid CDU Test Methodology and Performance Rating, 2024.
- Open Compute Project, Colocation Facility Guidelines for Deployment of OCP Racks.
- NVIDIA, "Hardware," DGX Grace Blackwell Rack Scale Systems User Guide, updated March 2026.
- NVIDIA, "FAQ," Mission Control Software Administration Guide, GB200 NVL72 designed rack power.
- NVIDIA, "System Hardware and Components," NVL72 AI Factory Enterprise Reference Architecture, GB300 NVL72 rack power.
- CoolIT Systems, "AHx240" liquid-to-air coolant distribution unit.
- NVIDIA, "Planning a Data Center Deployment," DGX SuperPOD Data Center Design Guide.
- NVIDIA, "Deploying the Bundles," DGX SuperPOD Cabling Data Centers Design Guide.