47 citations · 83 across the 9 of their papers we have counts for
9 papers
Knowledge-aware Alert Aggregation in Large-scale Cloud Systems: a Hybrid Approach
Jinxi Kuang, Jinyang Liu, Junjie Huang +6
Due to the scale and complexity of cloud systems, a system failure would trigger an "alert storm", i.e., massive correlated alerts. Although these alerts can be traced back to a fe…
FaultProfIT: Hierarchical Fault Profiling of Incident Tickets in Large-scale Cloud Systems
Junjie Huang, Jinyang Liu, Zhuangbin Chen +7
Postmortem analysis is essential in the management of incidents within cloud systems, which provides valuable insights to improve system's reliability and robustness. At CloudA, fa…
Go Static: Contextualized Logging Statement Generation
Yichen Li, Yintong Huo, Renyi Zhong +6
Logging practices have been extensively investigated to assist developers in writing appropriate logging statements for documenting software behaviors. Although numerous automatic…
MTAD: Tools and Benchmarks for Multivariate Time Series Anomaly Detection
Jinyang Liu, Wenwei Gu, Zhuangbin Chen +3
Key Performance Indicators (KPIs) are essential time-series metrics for ensuring the reliability and stability of many software systems. They faithfully record runtime states to fa…
A Roadmap towards Intelligent Operations for Reliable Cloud Computing Systems
Yintong Huo, Cheryl Lee, Jinyang Liu +2
The increasing complexity and usage of cloud systems have made it challenging for service providers to ensure reliability. This paper highlights two main challenges, namely interna…
Practical Anomaly Detection over Multivariate Monitoring Metrics for Online Services
Jinyang Liu, Tianyi Yang, Zhuangbin Chen +4
As modern software systems continue to grow in terms of complexity and volume, anomaly detection on multivariate monitoring metrics, which profile systems' health status, becomes m…