Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Detecting Silent Failures in Multi-Agentic AI Trajectories
Divya Pathak, Harshit Kumar, Anuska Roy +3
Multi-Agentic AI systems, powered by large language models (LLMs), are inherently non-deterministic and prone to silent failures such as drift, cycles, and missing details in outpu…
cs.AI2025
ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
Saurabh Jha, Rohan Arora, Yuji Watanabe +40
Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a…