collaborators
Showing cs.DCShow all

7 papers · 1 filter

cs.DC2026

Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads

Dominik Scheinert, Alexander Acker, Thorsten Wittkopp +6

Large language model (LLM) services have become an integral part of search, assistance, and decision-making applications. However, unlike traditional web or microservices, the hard…

cs.DC2026

Distributed LLM Pretraining During Renewable Curtailment Windows: A Feasibility Study

Philipp Wiesner, Soeren Becker, Brett Cornick +3

Training large language models (LLMs) requires substantial compute and energy. At the same time, renewable energy sources regularly produce more electricity than the grid can absor…

cs.DC2025

What happens when nanochat meets DiLoCo?

Alexander Acker, Soeren Becker, Sasho Nedelkoski +3

Although LLM training is typically centralized with high-bandwidth interconnects and large compute budgets, emerging methods target communication-constrained training in distribute…

cs.DC2025

Sizey: Memory-Efficient Execution of Scientific Workflow Tasks

Jonathan Bader, Fabian Skalski, Fabian Lehmann +4

As the amount of available data continues to grow in fields as diverse as bioinformatics, physics, and remote sensing, the importance of scientific workflows in the design and impl…

cs.DC2024

Privacy-Preserving Sharing of Data Analytics Runtime Metrics for Performance Modeling

Jonathan Will, Dominik Scheinert, Jan Bode +3

Performance modeling for large-scale data analytics workloads can improve the efficiency of cluster resource allocations and job scheduling. However, the performance of these workl…

cs.DC2024

Daedalus: Self-Adaptive Horizontal Autoscaling for Resource Efficiency of Distributed Stream Processing Systems

Benjamin J. J. Pfister, Dominik Scheinert, Morgan K. Geldenhuys +1

Distributed Stream Processing (DSP) systems are capable of processing large streams of unbounded data, offering high throughput and low latencies. To maintain a stable Quality of S…