From the 1 of 6 linked papers with an AI index.
3 citations · 3 across the 3 of their papers we have counts for
6 papers
Don't Predict, Prioritize: Rethinking GPU Reliability Assessment
Difeng Ma, Changhua Pei, Yuanwei Lu +7
The paper proposes HeaRank, a learning-to-rank framework that ranks GPU nodes by their relative failure risk instead of predicting exact failure times, showing improved detection o…
Agent System Operations: Categorization, Challenges, and Future Directions
Zexin Wang, Changhua Pei, Yuanhao Liu +10
As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional sys…
KAN-AD: Time Series Anomaly Detection with Kolmogorov-Arnold Networks
Quan Zhou, Changhua Pei, Fei Sun +6
Time series anomaly detection (TSAD) underpins real-time monitoring in cloud services and web systems, allowing rapid identification of anomalies to prevent costly failures. Most T…
ViTs: Teaching Machines to See Time Series Anomalies Like Human Experts
Zexin Wang, Changhua Pei, Yang Liu +8
Web service administrators must ensure the stability of multiple systems by promptly detecting anomalies in Key Performance Indicators (KPIs). Achieving the goal of "train once, in…
TShape: Rescuing Machine Learning Models from Complex Shapelet Anomalies
Hang Cui, Jingjing Li, Haotian Si +4
Time series anomaly detection (TSAD) is critical for maintaining the reliability of modern IT infrastructures, where complex anomalies frequently arise in highly dynamic environmen…
A Survey on AgentOps: Categorization, Challenges, and Future Directions
Zexin Wang, Jingjing Li, Quan Zhou +7
As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional sys…