Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation
Yeonjun In, Mehrab Tanjim, Jayakumar Subramanian +6
Failure attribution is essential for diagnosing and improving multi-agent systems (MAS), yet existing benchmarks and methods largely assume a single deterministic root cause for ea…
cs.AI2025
Evaluation and Incident Prevention in an Enterprise AI Assistant
Akash V. Maharaj, David Arbour, Daniel Lee +6
Enterprise AI Assistants are increasingly deployed in domains where accuracy is paramount, making each erroneous output a potentially significant incident. This paper presents a co…