Towards Advanced Monitoring for Scientific Workflows
arXiv:2211.12744 · doi:10.1109/BigData55660.2022.10020864
Abstract
Scientific workflows consist of thousands of highly parallelized tasks executed in a distributed environment involving many components. Automatic tracing and investigation of the components' and tasks' performance metrics, traces, and behavior are necessary to support the end user with a level of abstraction since the large amount of data cannot be analyzed manually. The execution and monitoring of scientific workflows involves many components, the cluster infrastructure, its resource manager, the workflow, and the workflow tasks. All components in such an execution environment access different monitoring metrics and provide metrics on different abstraction levels. The combination and analysis of observed metrics from different components and their interdependencies are still widely unregarded. We specify four different monitoring layers that can serve as an architectural blueprint for the monitoring responsibilities and the interactions of components in the scientific workflow execution context. We describe the different monitoring metrics subject to the four layers and how the layers interact. Finally, we examine five state-of-the-art scientific workflow management systems (SWMS) in order to assess which steps are needed to enable our four-layer-based approach.
Paper accepted in 2022 IEEE International Conference on Big Data Workshop SCDM 2022
References in corpus (8)
- Tarema: Adaptive Resource Allocation for Scalable Scientific Workflows in Heterogeneous Clusters
- Lotaru: Locally Estimating Runtimes of Scientific Workflow Tasks in Heterogeneous Clusters
- Reshi: Recommending Resources for Scientific Workflow Tasks on Heterogeneous Infrastructures
- Leveraging Reinforcement Learning for Task Resource Allocation in Scientific Workflows
- Get Your Memory Right: The Crispy Resource Allocation Assistant for Large-Scale Data Processing
- Macaw: The Machine Learning Magnetometer Calibration Workflow
- Ruya: Memory-Aware Iterative Optimization of Cluster Configurations for Big Data Processing
- Perona: Robust Infrastructure Fingerprinting for Resource-Efficient Big Data Analytics
Cited by in corpus (5)
- Leveraging Reinforcement Learning for Task Resource Allocation in Scientific Workflows
- KS+: Predicting Workflow Task Memory Usage Over Time
- The Common Workflow Scheduler Interface: Status Quo and Future Plans
- Scaling on Frontier: Uncertainty Quantification Workflow Applications using ExaWorks to Enable Full System Utilization
- Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications