works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.DC2026

Don't Predict, Prioritize: Rethinking GPU Reliability Assessment

Difeng Ma, Changhua Pei, Yuanwei Lu +7

The paper proposes HeaRank, a learning-to-rank framework that ranks GPU nodes by their relative failure risk instead of predicting exact failure times, showing improved detection o…

cs.SE2026

UModel: An Agent-Ready Observability Data Modeling Method at Scale

Changhua Pei, Zheyuan Li, Zexin Wang +10

When networked system failures occur, automatically performing Root Cause Analysis (RCA) using observability data is critical for ensuring networked system reliability. Recently, L…

cs.MA2026

Agent System Operations: Categorization, Challenges, and Future Directions

Zexin Wang, Changhua Pei, Yuanhao Liu +10

As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional sys…

cs.LG2025

ViTs: Teaching Machines to See Time Series Anomalies Like Human Experts

Zexin Wang, Changhua Pei, Yang Liu +8

Web service administrators must ensure the stability of multiple systems by promptly detecting anomalies in Key Performance Indicators (KPIs). Achieving the goal of "train once, in…

cs.SE2025

TShape: Rescuing Machine Learning Models from Complex Shapelet Anomalies

Hang Cui, Jingjing Li, Haotian Si +4

Time series anomaly detection (TSAD) is critical for maintaining the reliability of modern IT infrastructures, where complex anomalies frequently arise in highly dynamic environmen…

cs.AI2025

A Survey on AgentOps: Categorization, Challenges, and Future Directions

Zexin Wang, Jingjing Li, Quan Zhou +7

As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional sys…