Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Detect Before You Attribute: Cascade Failure Attribution for Multi-Agent Systems
Jiayi Zhang, Zexin Wang, Degang Sun +4
Large language model (LLM)-based agents have shown strong potential in solving complex tasks through multi-step reasoning, yet they remain vulnerable to execution failures. Accurat…
cs.AI2026
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
Yunfei Zhang, Boyu Feng, Changhua Pei +14
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then in…
cs.AI2025
A Survey on AgentOps: Categorization, Challenges, and Future Directions
Zexin Wang, Jingjing Li, Quan Zhou +7
As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional sys…