1 citations · 1 across the 9 of their papers we have counts for
3 papers · 1 filter
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
Chenkai Zhang, Yiran Li, Yifang Tian +2
Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, ye…
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios
Jackson Clark, Yiming Su, Saad Mohammad Rafid Pial +5
AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE). Current SRE benchmarks are limited to…
GALA: Can Graph-Augmented Large Language Model Agentic Workflows Elevate Root Cause Analysis?
Yifang Tian, Yaming Liu, Zichun Chong +2
Root cause analysis (RCA) in microservice systems is challenging, requiring on-call engineers to rapidly diagnose failures across heterogeneous telemetry such as metrics, logs, and…