1 citations · 1 across the 9 of their papers we have counts for
12 papers
Detect Before You Attribute: Cascade Failure Attribution for Multi-Agent Systems
Jiayi Zhang, Zexin Wang, Degang Sun +4
Large language model (LLM)-based agents have shown strong potential in solving complex tasks through multi-step reasoning, yet they remain vulnerable to execution failures. Accurat…
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
Yunfei Zhang, Boyu Feng, Changhua Pei +14
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then in…
Don't Predict, Prioritize: Rethinking GPU Reliability Assessment
Difeng Ma, Changhua Pei, Yuanwei Lu +7
The reliability of Graphics Processing Units (GPUs) is a criticalbottleneck for modern large-scale AI infrastructure, where a sin-gle node failure can disrupt synchronous training…
UModel: An Agent-Ready Observability Data Modeling Method at Scale
Changhua Pei, Zheyuan Li, Zexin Wang +10
When networked system failures occur, automatically performing Root Cause Analysis (RCA) using observability data is critical for ensuring networked system reliability. Recently, L…
Agent System Operations: Categorization, Challenges, and Future Directions
Zexin Wang, Changhua Pei, Yuanhao Liu +10
As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional sys…
KairosVL: Orchestrating Time Series and Semantics for Unified Reasoning
Haotian Si, Changhua Pei, Xiao He +9
Driven by the increasingly complex and decision-oriented demands of time series analysis, we introduce the Semantic-Conditional Time Series Reasoning task, which extends convention…