works on

From the 1 of 17 linked papers with an AI index.

collaborators

17 papers

cs.AI2026

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

Yunfei Zhang, Boyu Feng, Changhua Pei +14

When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then in…

cs.DC2026

Don't Predict, Prioritize: Rethinking GPU Reliability Assessment

Difeng Ma, Changhua Pei, Yuanwei Lu +7

The paper proposes HeaRank, a learning-to-rank framework that ranks GPU nodes by their relative failure risk instead of predicting exact failure times, showing improved detection o…

cs.CV2026

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding

Xianzhi Ma, Shujun Wang, Xiaohan Li +3

Ultra-High-Resolution (UHR) remote sensing image understanding requires Vision-Language Models (VLMs) to capture both the global scene layout and sparse yet task-critical local det…

cs.SE2026

A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

Yuanhong Cai, Xiaohui Nie, Kanglin Yin +8

LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data. However, existing…

cs.SE2026

UModel: An Agent-Ready Observability Data Modeling Method at Scale

Changhua Pei, Zheyuan Li, Zexin Wang +10

When networked system failures occur, automatically performing Root Cause Analysis (RCA) using observability data is critical for ensuring networked system reliability. Recently, L…

cs.MA2026

Agent System Operations: Categorization, Challenges, and Future Directions

Zexin Wang, Changhua Pei, Yuanhao Liu +10

As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional sys…