paper

Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems

arXiv:2603.00468

Abstract

LLM agents are increasingly explored for automating root cause analysis (RCA) in cloud-native systems, creating a need to evaluate both diagnostic correctness and the quality of the supporting investigation. Existing static benchmarks offer repeatable inputs but limited system-facing interaction, while live testbeds expose realistic tools but hinder controlled comparison because incident evidence varies across runs; both paradigms focus primarily on final answers. To address these limitations, we present Cloud-OpsBench, an evaluation infrastructure for interactive and evidence-grounded cloud RCA. It comprises 754 runtime-verified cases across 57 fault types on two microservice workloads spanning application services and Kubernetes platform layers. Each fault is captured as a state snapshot and replayed through standard diagnostic interfaces, with outcome labels and diagnostic evidence graphs that enable matched comparisons and process-level analysis. Across ten LLM agents, the strongest Joint RCA Accuracy (JRA) reaches 0.76 on OnlineBoutique and 0.68 on TrainTicket, while the corresponding Evidence Closure Rates (ECR) are only 0.38 and 0.15. This outcome--process gap shows that final-answer correctness alone substantially overestimates agents' ability to perform evidence-grounded diagnosis.

14 pages, 3 figures

Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems · wovepaper