5 papers · 1 filter
ErrorPrism: Reconstructing Error Propagation Paths in Cloud Service Systems
Junsong Pu, Yichen Li, Zhuangbin Chen +6
Reliability management in cloud service systems is challenging due to the cascading effect of failures. Error wrapping, a practice prevalent in modern microservice development, enr…
LogPilot: Intent-aware and Scalable Alert Diagnosis for Large-scale Online Service Systems
Zhihan Jiang, Jinyang Liu, Yichen Li +7
Effective alert diagnosis is essential for ensuring the reliability of large-scale online service systems. However, on-call engineers are often burdened with manually inspecting ma…
Adaptive and Efficient Log Parsing as a Cloud Service
Zeyan Li, Jie Song, Tieying Zhang +6
Logs are a critical data source for cloud systems, enabling advanced features like monitoring, alerting, and root cause analysis. However, the massive scale and diverse formats of…
TickIt: Leveraging Large Language Models for Automated Ticket Escalation
Fengrui Liu, Xiao He, Tieying Zhang +6
In large-scale cloud service systems, support tickets serve as a critical mechanism for resolving customer issues and maintaining service quality. However, traditional manual ticke…
Flow-of-Action: SOP Enhanced LLM-Based Multi-Agent System for Root Cause Analysis
Changhua Pei, Zexin Wang, Fengrui Liu +9
In the realm of microservices architecture, the occurrence of frequent incidents necessitates the employment of Root Cause Analysis (RCA) for swift issue resolution. It is common t…