9 papers
ORCA: Observability-Grounded Program Repair for Microservice Incidents
Yuanchen Gao, Yifang Tian, Yiran Li +2
Microservice failures are often diagnosed from operational telemetry. However, automated program repair systems usually start from issue reports, localized code context, or failing…
GALA: Graph-Augmented LLM Agents for Root Cause Analysis and Incident Response in Microservices
Yifang Tian, Yaming Liu, Zichun Chong +3
Microservice root cause analysis (RCA) requires correlating failures across heterogeneous telemetry within complex service dependency graphs. Existing methods often rely on a singl…
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
Chenkai Zhang, Yiran Li, Yifang Tian +2
Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, ye…
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios
Jackson Clark, Yiming Su, Saad Mohammad Rafid Pial +5
AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE). Current SRE benchmarks are limited to…
Epoch-based Optimistic Concurrency Control in Geo-replicated Databases
Yunhao Mao, Harunari Takata, Michail Bachras +4
Geo-distribution is essential for modern online applications to ensure service reliability and high availability. However, supporting high-performance serializable transactions in…
Ksurf-Drone: Attention Kalman Filter for Contextual Bandit Optimization in Cloud Resource Allocation
Michael Dang'ana, Yuqiu Zhang, Hans-Arno Jacobsen
Resource orchestration and configuration parameter search are key concerns for container-based infrastructure in cloud data centers. Large configuration search space and cloud unce…