paper

Data Center Life Cycle Co-Design Optimization

arXiv:2606.15408

Abstract

Liquid cooled supercomputers dissipate tens of megawatts of waste heat through cooling plants organized as parallel subloops that serve coolant distribution units. The number of subloops and the assignment of units to them are design decisions fixed at construction, yet they have not been systematically optimized at this scale. We present a framework that integrates operational energy from a validated control optimizer, embodied carbon and capital cost from a bill of materials, maintenance over the service life, and expected unplanned downtime from a component level reliability model. All 611 ways of partitioning the 25 coolant distribution units of the Frontier supercomputer into two through six subloops are evaluated. When redundancy is not costed, the optimum is two subloops holding 14 and 11 units, at 3,320.7 tonnes of carbon dioxide equivalent and 3,987k dollars over a 7 year horizon, saving 35.7 tonnes and 63k dollars compared to the documented as built configuration of three duty subloops holding 14, 6 and 5 units. The difference is driven by piping rather than by operational energy, and the identity of the optimum is unchanged across 15 sensitivity scenarios and both extremes of physical unit grouping, although its margin narrows under compact grouping. When the N+1 standby train that the plant actually carries is priced, the optimum moves to four or five duty subloops, adjacent to the three the plant runs and far from the unconstrained answer, and the semi-analytical decision rule reproduces this shift across four leadership class systems. Redundancy policy, not cost or carbon, is what sets the subloop count. A conversion of the built plant is shown not to pay back, so the framework is a greenfield design tool.

30 pages, 11 figures

Data Center Life Cycle Co-Design Optimization · wovepaper