collaborators

9 papers

cs.SE2026

ORCA: Observability-Grounded Program Repair for Microservice Incidents

Yuanchen Gao, Yifang Tian, Yiran Li +2

Microservice failures are often diagnosed from operational telemetry. However, automated program repair systems usually start from issue reports, localized code context, or failing…

cs.SE2026

GALA: Graph-Augmented LLM Agents for Root Cause Analysis and Incident Response in Microservices

Yifang Tian, Yaming Liu, Zichun Chong +3

Microservice root cause analysis (RCA) requires correlating failures across heterogeneous telemetry within complex service dependency graphs. Existing methods often rely on a singl…

cs.AI2026

When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry

Chenkai Zhang, Yiran Li, Yifang Tian +2

Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, ye…

cs.AI2026

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios

Jackson Clark, Yiming Su, Saad Mohammad Rafid Pial +5

AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE). Current SRE benchmarks are limited to…

cs.DB2026

Epoch-based Optimistic Concurrency Control in Geo-replicated Databases

Yunhao Mao, Harunari Takata, Michail Bachras +4

Geo-distribution is essential for modern online applications to ensure service reliability and high availability. However, supporting high-performance serializable transactions in…

cs.DC2025

Ksurf-Drone: Attention Kalman Filter for Contextual Bandit Optimization in Cloud Resource Allocation

Michael Dang'ana, Yuqiu Zhang, Hans-Arno Jacobsen

Resource orchestration and configuration parameter search are key concerns for container-based infrastructure in cloud data centers. Large configuration search space and cloud unce…