How Bad Can a Bug Get? An Empirical Analysis of Software Failures in the OpenStack Cloud Computing Platform
arXiv:1907.04055 · doi:10.1145/3338906.3338916
Abstract
Cloud management systems provide abstractions and APIs for programmatically configuring cloud infrastructures. Unfortunately, residual software bugs in these systems can potentially lead to high-severity failures, such as prolonged outages and data losses. In this paper, we investigate the impact of failures in the context widespread OpenStack cloud management system, by performing fault injection and by analyzing the impact of the resulting failures in terms of fail-stop behavior, failure detection through logging, and failure propagation across components. The analysis points out that most of the failures are not timely detected and notified; moreover, many of these failures can silently propagate over time and through components of the cloud management system, which call for more thorough run-time checks and fault containment.
12 pages, ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE '19)
Cited by in corpus (12)
- A Survey on Automated Log Analysis for Reliability Engineering
- Fault Injection Analytics: A Novel Approach to Discover Failure Modes in Cloud-Computing Systems
- ProFIPy: Programmable Software Fault Injection as-a-Service
- Mutiny! How does Kubernetes fail, and what can we do about it?
- Systematic Evaluation of Deep Learning Models for Log-based Failure Prediction
- A Comprehensive Study of Machine Learning Techniques for Log-Based Anomaly Detection
- Run-time Failure Detection via Non-intrusive Event Analysis in a Large-Scale Cloud Computing Platform
- The HitchHiker's Guide to High-Assurance System Observability Protection with Efficient Permission Switches
- A Comprehensive Survey of Logging in Software: From Logging Statements Automation to Log Mining and Analysis
- Towards Runtime Verification via Event Stream Processing in Cloud Computing Infrastructures
- Enhancing the Analysis of Software Failures in Cloud Computing Systems with Deep Learning
- Security-by-Design at the Telco Edge with OSS: Challenges and Lessons Learned