17 citations · 63 across the 11 of their papers we have counts for
11 papers
KS+: Predicting Workflow Task Memory Usage Over Time
Jonathan Bader, Ansgar Lößer, Lauritz Thamsen +2
Scientific workflow management systems enable the reproducible execution of data analysis pipelines on cluster infrastructures managed by resource managers such as Kubernetes, Slur…
Privacy-Preserving Sharing of Data Analytics Runtime Metrics for Performance Modeling
Jonathan Will, Dominik Scheinert, Jan Bode +3
Performance modeling for large-scale data analytics workloads can improve the efficiency of cluster resource allocations and job scheduling. However, the performance of these workl…
Demeter: Resource-Efficient Distributed Stream Processing under Dynamic Loads with Multi-Configuration Optimization
Morgan Geldenhuys, Dominik Scheinert, Odej Kao +1
Distributed Stream Processing (DSP) focuses on the near real-time processing of large streams of unbounded data. To increase processing capacities, DSP systems are able to dynamica…
Lotaru: Locally Predicting Workflow Task Runtimes for Resource Management on Heterogeneous Infrastructures
Jonathan Bader, Fabian Lehmann, Lauritz Thamsen +2
Many resource management techniques for task scheduling, energy and carbon efficiency, and cost optimization in workflows rely on a-priori task runtime knowledge. Building runtime…
Karasu: A Collaborative Approach to Efficient Cluster Configuration for Big Data Analytics
Dominik Scheinert, Philipp Wiesner, Thorsten Wittkopp +3
Selecting the right resources for big data analytics jobs is hard because of the wide variety of configuration options like machine type and cluster size. As poor choices can have…
Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?
Jonathan Will, Lauritz Thamsen, Dominik Scheinert +1
Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate…