3 citations · 3 across the 5 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
Guangba Yu, Genting Mai, Rui Wang +4
Alerts are critical for detecting anomalies in large-scale cloud systems, ensuring reliability and user experience. However, current systems generate overwhelming volumes of alerts…
cs.DC2025
eACGM: Non-instrumented Performance Tracing and Anomaly Detection towards Machine Learning Systems
Ruilin Xu, Zongxuan Xie, Pengfei Chen
We present eACGM, a full-stack AI/ML system monitoring framework based on eBPF. eACGM collects real-time performance data from key hardware components, including the GPU and networ…