From the 1 of 18 linked papers with an AI index.
18 papers
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
Yunfei Zhang, Boyu Feng, Changhua Pei +14
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then in…
Don't Predict, Prioritize: Rethinking GPU Reliability Assessment
Difeng Ma, Changhua Pei, Yuanwei Lu +7
The paper proposes HeaRank, a learning-to-rank framework that ranks GPU nodes by their relative failure risk instead of predicting exact failure times, showing improved detection o…
UModel: An Agent-Ready Observability Data Modeling Method at Scale
Changhua Pei, Zheyuan Li, Zexin Wang +10
When networked system failures occur, automatically performing Root Cause Analysis (RCA) using observability data is critical for ensuring networked system reliability. Recently, L…
Agent System Operations: Categorization, Challenges, and Future Directions
Zexin Wang, Changhua Pei, Yuanhao Liu +10
As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional sys…
KAN-AD: Time Series Anomaly Detection with Kolmogorov-Arnold Networks
Quan Zhou, Changhua Pei, Fei Sun +6
Time series anomaly detection (TSAD) underpins real-time monitoring in cloud services and web systems, allowing rapid identification of anomalies to prevent costly failures. Most T…
From Time Series Analysis to Question Answering: A Survey in the LLM Era
Wei Li, Zhe Xie, Yuxuan Liang +4
Recently, Large Language Models (LLMs) have introduced a novel paradigm in Time Series Analysis (TSA), leveraging strong language capabilities to support tasks such as forecasting…