activity
20212024
most citedLotaru: Locally Predicting Workflow Task Runtimes for Resource Management on Heterogeneous Infrastructures

17 citations · 63 across the 11 of their papers we have counts for

collaborators

11 papers

cs.DC20245 cited

KS+: Predicting Workflow Task Memory Usage Over Time

Jonathan Bader, Ansgar Lößer, Lauritz Thamsen +2

Scientific workflow management systems enable the reproducible execution of data analysis pipelines on cluster infrastructures managed by resource managers such as Kubernetes, Slur…

cs.DC2024

Privacy-Preserving Sharing of Data Analytics Runtime Metrics for Performance Modeling

Jonathan Will, Dominik Scheinert, Jan Bode +3

Performance modeling for large-scale data analytics workloads can improve the efficiency of cluster resource allocations and job scheduling. However, the performance of these workl…

cs.DC2024

Demeter: Resource-Efficient Distributed Stream Processing under Dynamic Loads with Multi-Configuration Optimization

Morgan Geldenhuys, Dominik Scheinert, Odej Kao +1

Distributed Stream Processing (DSP) focuses on the near real-time processing of large streams of unbounded data. To increase processing capacities, DSP systems are able to dynamica…

cs.DC202317 cited

Lotaru: Locally Predicting Workflow Task Runtimes for Resource Management on Heterogeneous Infrastructures

Jonathan Bader, Fabian Lehmann, Lauritz Thamsen +2

Many resource management techniques for task scheduling, energy and carbon efficiency, and cost optimization in workflows rely on a-priori task runtime knowledge. Building runtime…

cs.DC20234 cited

Karasu: A Collaborative Approach to Efficient Cluster Configuration for Big Data Analytics

Dominik Scheinert, Philipp Wiesner, Thorsten Wittkopp +3

Selecting the right resources for big data analytics jobs is hard because of the wide variety of configuration options like machine type and cluster size. As poor choices can have…

cs.DC20232 cited

Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?

Jonathan Will, Lauritz Thamsen, Dominik Scheinert +1

Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate…