From the 1 of 14 linked papers with an AI index.
9 papers · 1 filter
AgentTether: Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operation
Chenyu Zhao, Shenglin Zhang, Wenwei Gu +5
Large language model (LLM) agents are increasingly used for multi-step, stateful tool-use tasks, yet production reliability remains limited. Unlike static software repair, agent re…
KRCA: An Efficient Root Cause Analysis System in Hyper-scale Microservice Systems via Agentic AI
Jiamin Jiang, Jingfei Feng, Yu Luo +10
Hyper-scale microservice systems have become the standard infrastructure for large-scale Internet companies. These systems consist of numerous loosely coupled microservices that ev…
A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis
Yuanhong Cai, Xiaohui Nie, Kanglin Yin +8
LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data. However, existing…
Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents
Chenyu Zhao, Shenglin Zhang, Yihang Lin +7
Software engineering agents are increasingly deployed in evaluable engineering environments, yet post-failure recovery remains costly, manual, and ad hoc. Existing systems expose t…
EvidenT: An Evidence-Preserving Framework for Iterative System-Level Package Repair
Chenyu Zhao, Minghua Ma, Shenglin Zhang +5
Frequent toolchain updates and growing ISA diversity have made system-level software package repair increasingly important. Diagnosing and repairing build failures remains challeng…
Which Types of Heterogeneity Matter for Root Cause Localization in Microservice Systems ?
Runzhou Wang, Shenglin Zhang, Wenwei Gu +5
Microservice root cause localization is fundamentally challenged by the inherent heterogeneity of cloud-native systems, which encompasses diverse observability data and multiple sy…