5 papers
Designing Datacenter Power Delivery Hierarchies for the AI Era
Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare +2
Demand for AI accelerators is rapidly increasing rack power density, with projections approaching 1MW per deployment by 2027. This poses a major challenge for datacenter power deli…
EasyRider: Mitigating Power Transients in Datacenter-Scale Training Workloads
Dillon Jensen, Obi Nnorom, Grant Wilkins +4
Large-scale AI model training workloads use thousands of GPUs operating in tightly synchronized loops. During synchronous communication, start-up, shut-down, and checkpointing, GPU…
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
Grant Wilkins, Fiodar Kazhamiaka, Ram Rajagopal
Datacenter operators and electrical utilities rely on power traces at different spatiotemporal scales. Operators use fine-grained traces for provisioning, facility management, and…
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
Grant Wilkins, Srinivasan Keshav, Richard Mortier
The rapid adoption of large language models (LLMs) has led to significant advances in natural language processing and text generation. However, the energy consumed through LLM mode…
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
Grant Wilkins, Srinivasan Keshav, Richard Mortier
Both the training and use of Large Language Models (LLMs) require large amounts of energy. Their increasing popularity, therefore, raises critical concerns regarding the energy eff…