From the 1 of 18 linked papers with an AI index.
3 citations · 3 across the 6 of their papers we have counts for
4 papers · 1 filter
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
Yunfei Zhang, Boyu Feng, Changhua Pei +14
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then in…
KairosVL: Orchestrating Time Series and Semantics for Unified Reasoning
Haotian Si, Changhua Pei, Xiao He +9
Driven by the increasingly complex and decision-oriented demands of time series analysis, we introduce the Semantic-Conditional Time Series Reasoning task, which extends convention…
A Survey on AgentOps: Categorization, Challenges, and Future Directions
Zexin Wang, Jingjing Li, Quan Zhou +7
As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional sys…
OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models
Yuhe Liu, Changhua Pei, Longlong Xu +13
Information Technology (IT) Operations (Ops), particularly Artificial Intelligence for IT Operations (AIOps), is the guarantee for maintaining the orderly and stable operation of e…