10 citations · 16 across the 7 of their papers we have counts for
7 papers
Privacy-Preserving Sharing of Data Analytics Runtime Metrics for Performance Modeling
Jonathan Will, Dominik Scheinert, Jan Bode +3
Performance modeling for large-scale data analytics workloads can improve the efficiency of cluster resource allocations and job scheduling. However, the performance of these workl…
Demeter: Resource-Efficient Distributed Stream Processing under Dynamic Loads with Multi-Configuration Optimization
Morgan Geldenhuys, Dominik Scheinert, Odej Kao +1
Distributed Stream Processing (DSP) focuses on the near real-time processing of large streams of unbounded data. To increase processing capacities, DSP systems are able to dynamica…
Karasu: A Collaborative Approach to Efficient Cluster Configuration for Big Data Analytics
Dominik Scheinert, Philipp Wiesner, Thorsten Wittkopp +3
Selecting the right resources for big data analytics jobs is hard because of the wide variety of configuration options like machine type and cluster size. As poor choices can have…
Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?
Jonathan Will, Lauritz Thamsen, Dominik Scheinert +1
Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate…
PULL: Reactive Log Anomaly Detection Based On Iterative PU Learning
Thorsten Wittkopp, Dominik Scheinert, Philipp Wiesner +2
Due to the complexity of modern IT services, failures can be manifold, occur at any stage, and are hard to detect. For this reason, anomaly detection applied to monitoring data suc…
Reshi: Recommending Resources for Scientific Workflow Tasks on Heterogeneous Infrastructures
Jonathan Bader, Fabian Lehmann, Alexander Groth +5
Scientific workflows typically comprise a multitude of different processing steps which often are executed in parallel on different partitions of the input data. These executions,…