3 papers
cs.SE2026
A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis
Yuanhong Cai, Xiaohui Nie, Kanglin Yin +8
LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data. However, existing…
cs.SE2026
UModel: An Agent-Ready Observability Data Modeling Method at Scale
Changhua Pei, Zheyuan Li, Zexin Wang +10
When networked system failures occur, automatically performing Root Cause Analysis (RCA) using observability data is critical for ensuring networked system reliability. Recently, L…
cs.LG2026
OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning
Jingkai He, Pengfei Chen, Chenghui Wu +7
In the field of software operations, Large Language Models (LLMs) have attracted increasing attention. However, existing research has not yet achieved efficient and effective endto…