most citedWhich Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

cs.AI2026

Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering

Wenbo Zhang, Zhongxiang Sun, Zhiguang Han +1

Sparse autoencoder (SAE)-based steering has been widely used to address knowledge conflicts by guiding LLMs to be more faithful to the contextual knowledge. Existing methods usuall…

cs.AI2026

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

Kunlun Zhu, Xuyan Ye, Zhiguang Han +9

LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but pro…

cs.AI2026

Harnessing Agentic Evolution

Jiayi Zhang, Yongfeng Gu, Jianhao Ruan +10

Agentic evolution has emerged as a powerful paradigm for improving programs, workflows, and scientific solutions by iteratively generating candidates, evaluating them, and using fe…

cs.AI2026

ACE-Router: Generalizing History-Aware Routing from MCP Tools to the Agent Web

Zhiyuan Yao, Zishan Xu, Yifu Guo +6

With the rise of the Agent Web and Model Context Protocol (MCP), the agent ecosystem is evolving into an open collaborative network, exponentially increasing accessible tools. Howe…

cs.CL2025

MultiJustice: A Chinese Dataset for Multi-Party, Multi-Charge Legal Prediction

Xiao Wang, Jiahuan Pei, Diancheng Shui +4

Legal judgment prediction offers a compelling method to aid legal practitioners and researchers. However, the research question remains relatively under-explored: Should multiple d…

cs.MA2025★ 1 cited

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Shaokun Zhang, Ming Yin, Jieyu Zhang +8

Failure attribution in LLM multi-agent systems-identifying the agent and step responsible for task failures-provides crucial clues for systems debugging but remains underexplored a…