collaborators

6 papers

cs.DC2026

Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads

Dominik Scheinert, Alexander Acker, Thorsten Wittkopp +6

Large language model (LLM) services have become an integral part of search, assistance, and decision-making applications. However, unlike traditional web or microservices, the hard…

cs.DC2026

Distributed LLM Pretraining During Renewable Curtailment Windows: A Feasibility Study

Philipp Wiesner, Soeren Becker, Brett Cornick +3

Training large language models (LLMs) requires substantial compute and energy. At the same time, renewable energy sources regularly produce more electricity than the grid can absor…

cs.DC2025

What happens when nanochat meets DiLoCo?

Alexander Acker, Soeren Becker, Sasho Nedelkoski +3

Although LLM training is typically centralized with high-bandwidth interconnects and large compute budgets, emerging methods target communication-constrained training in distribute…

cs.LG2025

Distributed Low-Communication Training with Decoupled Momentum Optimization

Sasho Nedelkoski, Alexander Acker, Odej Kao +2

The training of large models demands substantial computational resources, typically available only in data centers with high-bandwidth interconnects. However, reducing the reliance…

cs.DC2025

Predicting the Performance of Scientific Workflow Tasks for Cluster Resource Management: An Overview of the State of the Art

Jonathan Bader, Kathleen West, Soeren Becker +5

Scientific workflow management systems support large-scale data analysis on cluster infrastructures. For this, they interact with resource managers which schedule workflow tasks on…

cs.DC2025

WOW: Workflow-Aware Data Movement and Task Scheduling for Dynamic Scientific Workflows

Fabian Lehmann, Jonathan Bader, Friedrich Tschirpke +6

Scientific workflows process extensive data sets over clusters of independent nodes, which requires a complex stack of infrastructure components, especially a resource manager (RM)…