4 papers
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
Guangba Yu, Genting Mai, Rui Wang +4
Alerts are critical for detecting anomalies in large-scale cloud systems, ensuring reliability and user experience. However, current systems generate overwhelming volumes of alerts…
eACGM: Non-instrumented Performance Tracing and Anomaly Detection towards Machine Learning Systems
Ruilin Xu, Zongxuan Xie, Pengfei Chen
We present eACGM, a full-stack AI/ML system monitoring framework based on eBPF. eACGM collects real-time performance data from key hardware components, including the GPU and networ…
FaaSRCA: Full Lifecycle Root Cause Analysis for Serverless Applications
Jin Huang, Pengfei Chen, Guangba Yu +3
Serverless becomes popular as a novel computing paradigms for cloud native services. However, the complexity and dynamic nature of serverless applications present significant chall…
Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis
Haiyu Huang, Cheng Chen, Kunyi Chen +6
Distributed traces contain valuable information but are often massive in volume, posing a core challenge in tracing framework design: balancing the tradeoff between preserving esse…