On the Potential of Execution Traces for Batch Processing Workload Optimization in Public Clouds
arXiv:2111.08759 · doi:10.1109/BigData52589.2021.9671275
Abstract
With the growing amount of data, data processing workloads and the management of their resource usage becomes increasingly important. Since managing a dedicated infrastructure is in many situations infeasible or uneconomical, users progressively execute their respective workloads in the cloud. As the configuration of workloads and resources is often challenging, various methods have been proposed that either quickly profile towards a good configuration or determine one based on data from previous runs. Still, performance data to train such methods is often lacking and must be costly collected. In this paper, we propose a collaborative approach for sharing anonymized workload execution traces among users, mining them for general patterns, and exploiting clusters of historical workloads for future optimizations. We evaluate our prototype implementation for mining workload execution graphs on a publicly available trace dataset and demonstrate the predictive value of workload clusters determined through traces only.
6 pages, 5 figures, 1 table
References in corpus (3)
Cited by in corpus (4)
- Get Your Memory Right: The Crispy Resource Allocation Assistant for Large-Scale Data Processing
- Training Data Reduction for Performance Models of Data Analytics Jobs in the Cloud
- Ruya: Memory-Aware Iterative Optimization of Cluster Configurations for Big Data Processing
- Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications