10 papers
Learning Process Energy Profiles from Node-Level Power Data
Jonathan Bader, Julius Irion, Jannis Kappel +4
The growing demand for data center capacity, driven by the growth of high-performance computing, cloud computing, and especially artificial intelligence, has led to a sharp increas…
Optimizing Memory Allocation in Distributed Clusters with Predictive Modeling
Jonathan Bader, Edgar Blumenthal, Marten Eckardt +5
In modern distributed systems, efficient resource allocation is a vital aspect to maintain scalability, reduce operational costs, and ensure fast execution even across heterogeneou…
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
Dominik Scheinert, Alexander Acker, Thorsten Wittkopp +6
Large language model (LLM) services have become an integral part of search, assistance, and decision-making applications. However, unlike traditional web or microservices, the hard…
What happens when nanochat meets DiLoCo?
Alexander Acker, Soeren Becker, Sasho Nedelkoski +3
Although LLM training is typically centralized with high-bandwidth interconnects and large compute budgets, emerging methods target communication-constrained training in distribute…
Distributed Low-Communication Training with Decoupled Momentum Optimization
Sasho Nedelkoski, Alexander Acker, Odej Kao +2
The training of large models demands substantial computational resources, typically available only in data centers with high-bandwidth interconnects. However, reducing the reliance…
Optimizing Microgrid Composition for Sustainable Data Centers
Julius Irion, Philipp Wiesner, Jonathan Bader +1
As computing energy demand continues to grow and electrical grid infrastructure struggles to keep pace, an increasing number of data centers are being planned with colocated microg…