A conventional data centre infrastructure management deployment answers three questions. How much power is the room drawing. How much cooling capacity is left. Which cabinet has space in it. Those answers are sufficient when a rack draws 5 to 10kW and the load is broadly flat across a working week.
A GPU cluster breaks all three. Researchers measuring an 8-GPU NVIDIA H100 HGX node during ResNet and Llama2-13b training recorded a maximum instantaneous draw of about 8.4kW against a manufacturer rating of 10.2kW, with the GPUs near full utilisation. That is one node. Put four in a cabinet and you are past 30kW, with the load moving in step with a job scheduler the facility team does not operate and often cannot see.
DCIM for AI GPU clusters is not the same product with bigger numbers in it. Three things change: how often you sample power, where you place thermal sensors, and what a unit of capacity means.
Power telemetry: the sampling interval is the argument
Most DCIM installations poll intelligent PDUs over SNMP every five minutes and store an average. On a database floor that is fine, because the load barely moves between samples. On a GPU rack it is close to useless, because the events that matter happen between polls.
Meta reported that during Llama 3 training, tens of thousands of GPUs can raise or lower power consumption at the same moment, for example when all of them wait for checkpointing or a collective communication to finish, or when a job starts or stops. The result was instant fluctuations of power consumption across the data centre on the order of tens of megawatts. At single-rack scale that is the same physics: a synchronised step across 32 GPUs in a cabinet moves tens of kilowatts in under a second.
Breakers respond to peaks. Averages do not contain peaks. So the useful telemetry configuration is per-outlet or per-branch metering sampled at seconds rather than minutes, storing both the peak within the interval and the time-weighted mean. You need the peak to size protection and the mean to size the utility feed and the cooling.
The same problem shows up in capacity forecasting. Plan cabinet loading on nameplate and you strand roughly 18% of your power on paper, based on the H100 node measurements above. Plan on five-minute averages and you undercount the transient and trip on a checkpoint boundary. Both numbers belong in the model, measured on your hardware and your workload rather than inherited from a datasheet.
Thermal monitoring moves from the room to the cold plate
Room-level temperature and humidity sensing was designed for air-cooled halls with hot and cold aisle containment. It tells you very little about a direct liquid cooling loop.
ASHRAE’s thermal guidelines classify facility supply water by its maximum allowable temperature, running W17, W27, W32, W40, W45 and W+ above 45 degrees. A liquid-cooled deployment is engineered to sit inside one of those classes, which means your DCIM has to instrument the loop against it: supply temperature, return temperature, the delta between them, flow rate per manifold, coolant distribution unit pressure differential, filter state and leak detection at rack and row level. A cabinet at 60kW rejecting heat to water has no thermal buffer. Room air temperature will not warn you in time.
The second thermal input is the one facility teams usually do not have access to. NVIDIA’s Data Center GPU Manager exposes per-device fields including GPU temperature, memory temperature, SM clock frequency and a clock-events bitmask that records why clocks were reduced, including software thermal slowdown. Exported to a time-series store, throttle minutes per GPU per day is a facility metric. It rises before a loop failure becomes an outage.
Meta’s Llama 3 405B run showed a diurnal 1 to 2% throughput variation caused by higher mid-day temperatures affecting GPU dynamic voltage and frequency scaling. That is a facility condition that only became visible in compute telemetry. If your DCIM and your GPU monitoring are separate systems with no shared clock, you will never see the correlation.
Capacity planning when a rack is 30 to 100kW
On a conventional floor, capacity is counted in rack units and a diversity factor. Racks peak at different times, so you plan the room against an aggregate that is comfortably below the sum of the nameplates.
Diversity is the assumption that fails first with AI workloads. A distributed training job is synchronous by design. Every GPU in the job steps together, waits together and checkpoints together. Across a pod, the diversity factor tends towards one. Capacity planning has to be done against coincident peak within each power block, not against a floor-wide average.
Three other constraints belong in the same model and rarely appear in legacy DCIM:
Floor loading. Dense GPU cabinets with liquid manifolds and busbars are heavy. Capacity is kilograms per square metre as well as kilowatts per cabinet.
Coolant capacity. A CDU has a rated heat rejection in kilowatts and a flow ceiling in litres per minute. Adding a rack to a loop that is already near its flow limit adds power draw you cannot cool.
Failure-domain headroom. If the design is concurrently maintainable, capacity in a block is what remains when one distribution path is down for maintenance, not what the block can carry with everything running.
Add to that a workload-side reserve. A training run that loses power mid-step restarts from its last checkpoint, so the cost of a capacity miscalculation is measured in GPU-hours, not in a failed request.
Where facility telemetry and compute telemetry have to meet
The two halves of the picture speak different protocols. The facility side runs SNMP, Modbus and BACnet against PDUs, UPS plant, CDUs and building management systems. The IT side runs Redfish against baseboard management controllers and DCGM against the GPUs, typically scraped into a time-series database.
The work is normalisation and correlation, not collection. A DCIM stack that is useful for GPU estates ingests both, aligns them on a common timestamp, and computes values neither side produces alone: rack headroom against breaker rating, coolant delta T against design, throttle events attributed to a cabinet, and energy per training step or per million tokens served.
That last one is worth building deliberately. The same H100 measurements found that raising the ResNet batch size from 512 to 4096 cut total training energy by a factor of four with the model architecture unchanged. Energy per unit of useful work is a tuneable number, and you cannot tune what you do not attribute.
Metering granularity also decides what you can report externally. NABERS rates Australian data centres on three streams: IT equipment, infrastructure, and whole facility, with the whole-facility tool intended for cases where internal metering does not permit the other two to be separated. Which rating you can claim is a function of how your metering was designed.
What to ask before you commit a cluster to a facility
Ask what the metering point is and how often it is sampled. Per-cabinet at five minutes is not the same as per-branch at one second, and only one of those will explain a trip.
Ask whether the colocation operator can give you facility telemetry as a feed, not just a portal. If you cannot correlate the coolant loop against your own GPU telemetry, you are running two blind systems.
Ask which liquid cooling temperature class the hall is engineered to, and what that does to your rack density budget at the warm end of it.
Ask what the coincident-peak assumption is for your power block, and whether it was set for mixed enterprise load or for synchronous GPU jobs.
Ask what the alarm thresholds are on coolant flow and leak detection, and who receives them at 3am.
None of this is exotic. It is the discipline that already applies to power and cooling in a well-run facility, moved down two levels of granularity and out one protocol boundary, so that the people running the cluster and the people running the building read the same numbers.
Related reading: Inside an AI-ready data centre: what makes it different and Choosing an AI infrastructure partner.
Frequently asked questions
What is DCIM and how is it different for AI workloads? Data centre infrastructure management is the software layer that monitors and plans physical facility resources: power, cooling, space and connectivity. For AI GPU clusters the difference is granularity and speed. Conventional DCIM samples at minutes and plans in rack units; GPU estates need sub-minute power sampling, coolant-loop instrumentation and capacity planning expressed in kilowatts and litres per minute per cabinet.
How often should power be sampled on a GPU rack? Fast enough to capture the transient that trips protection. Synchronised training jobs change power state across a whole pod within a second, so per-branch metering at second-level intervals, retaining both interval peak and mean, is the practical baseline. Five-minute averages hide exactly the event you are trying to size for.
Can existing DCIM software be extended to monitor GPUs? Usually yes, if it can ingest a second telemetry source and align it on time. GPU-level data comes from NVIDIA’s Data Center GPU Manager and from Redfish on the baseboard management controller, while the facility side stays on SNMP, Modbus and BACnet. The integration work is normalising both into one time-series view and computing derived metrics such as rack headroom and throttle minutes.
Why does GPU throttling matter to the facility team? Because it is often the earliest signal that cooling is drifting. A rising count of thermal clock-event reasons on GPUs in one cabinet points at a specific loop, manifold or filter well before room-level sensors register a change, and it also quantifies the cost, since throttled clocks mean the job is running slower on the same power.