4 papers
Detecting Silent Failures in Multi-Agentic AI Trajectories
Divya Pathak, Harshit Kumar, Anuska Roy +3
Multi-Agentic AI systems, powered by large language models (LLMs), are inherently non-deterministic and prone to silent failures such as drift, cycles, and missing details in outpu…
Unsupervised Cycle Detection in Agentic Applications
Felix George, Harshit Kumar, Divya Pathak +3
Agentic applications powered by Large Language Models exhibit non-deterministic behaviors that can form hidden execution cycles, silently consuming resources without triggering exp…
Metric Criticality Identification for Cloud Microservices
Akanksha Singal, Divya Pathak, Kaustabha Ray +3
Modern cloud-native applications built on microservice architectures present unprecedented challenges for system monitoring and alerting. Site Reliability Engineers (SREs) face the…
ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
Saurabh Jha, Rohan Arora, Yuji Watanabe +40
Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a…