3 papers
cs.DC2025
ClusterRCA: An End-to-End Approach for Network Fault Localization and Classification for HPC System
Yongqian Sun, Xijie Pan, Xiao Xiong +6
Network failure diagnosis is challenging yet critical for high-performance computing (HPC) systems. Existing methods cannot be directly applied to HPC scenarios due to data heterog…
cs.SE2024
Failure Diagnosis in Microservice Systems: A Comprehensive Survey and Analysis
Shenglin Zhang, Sibo Xia, Wenzhao Fan +6
Widely adopted for their scalability and flexibility, modern microservice systems present unique failure diagnosis challenges due to their independent deployment and dynamic intera…
cs.SE2023
Robust Multimodal Failure Detection for Microservice Systems
Chenyu Zhao, Minghua Ma, Zhenyu Zhong +10
Proactive failure detection of instances is vitally essential to microservice systems because an instance failure can propagate to the whole system and degrade the system's perform…