3 citations · 4 across the 10 of their papers we have counts for
6 papers · 1 filter
Detect Before You Attribute: Cascade Failure Attribution for Multi-Agent Systems
Jiayi Zhang, Zexin Wang, Degang Sun +4
Large language model (LLM)-based agents have shown strong potential in solving complex tasks through multi-step reasoning, yet they remain vulnerable to execution failures. Accurat…
LongRCA Bench: Root-Cause Localization in Long-Horizon Agent Trajectories
Yunfei Zhang, Boyu Feng, Changhua Pei +14
In long agent executions, an early error can persist through later actions and checks, while evidence needed to trace its origin is dispersed across the history. Short histories of…
Don't Predict, Prioritize: Rethinking GPU Reliability Assessment
Difeng Ma, Changhua Pei, Yuanwei Lu +7
The reliability of Graphics Processing Units (GPUs) is a criticalbottleneck for modern large-scale AI infrastructure, where a sin-gle node failure can disrupt synchronous training…
UModel: An Agent-Ready Observability Data Modeling Method at Scale
Changhua Pei, Zheyuan Li, Zexin Wang +10
When networked system failures occur, automatically performing Root Cause Analysis (RCA) using observability data is critical for ensuring networked system reliability. Recently, L…
Agent System Operations: Categorization, Challenges, and Future Directions
Zexin Wang, Changhua Pei, Yuanhao Liu +10
As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional sys…
KairosVL: Orchestrating Time Series and Semantics for Unified Reasoning
Haotian Si, Changhua Pei, Xiao He +9
Driven by the increasingly complex and decision-oriented demands of time series analysis, we introduce the Semantic-Conditional Time Series Reasoning task, which extends convention…