7 papers
A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis
Yuanhong Cai, Xiaohui Nie, Kanglin Yin +8
LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data. However, existing…
UModel: An Agent-Ready Observability Data Modeling Method at Scale
Changhua Pei, Zheyuan Li, Zexin Wang +10
When networked system failures occur, automatically performing Root Cause Analysis (RCA) using observability data is critical for ensuring networked system reliability. Recently, L…
Graph of States: Solving Abductive Tasks with Large Language Models
Yu Luo, Rongchen Gao, Lu Teng +9
Logical reasoning encompasses deduction, induction, and abduction. However, while Large Language Models (LLMs) have effectively mastered the former two, abductive reasoning remains…
OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning
Jingkai He, Pengfei Chen, Chenghui Wu +8
In the field of software operations, Large Language Models (LLMs) have attracted increasing attention. However, existing research has not yet achieved efficient and effective endto…
OpsAgent: An Evolving Multi-agent System for Incident Management in Microservices
Yu Luo, Jiamin Jiang, Jingfei Feng +6
Incident management (IM) is central to the reliability of large-scale microservice systems. Yet manual IM, where on-call engineers examine metrics, logs, and traces is labor-intens…
TrioXpert: An Automated Incident Management Framework for Microservice System
Yongqian Sun, Yu Luo, Xidao Wen +5
Automated incident management plays a pivotal role in large-scale microservice systems. However, many existing methods rely solely on single-modal data (e.g., metrics, logs, and tr…