works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.AI2026

Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

Gen Dong, Yanjie Gao, Liqun Li +3

Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving…

cs.MA2026

Towards a Systems Foundation for Agentic Cloud Management

Minghao Li, Ziqian Liu, Ziyu Mao +5

The paper proposes CloudWeaver, a systems layer that lets autonomous agents safely manage cloud resources by providing scoped views and coordinating concurrent operations, ensuring…

cs.AI2026

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios

Jackson Clark, Yiming Su, Saad Mohammad Rafid Pial +5

AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE). Current SRE benchmarks are limited to…

cs.DC2026

STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds

Yinfang Chen, Jiaqi Pan, Jackson Clark +7

In cloud-scale systems, failures are the norm. A distributed computing cluster exhibits hundreds of machine failures and thousands of disk failures; software bugs and misconfigurat…

cs.DC2025

An Empirical Study of Production Incidents in Generative AI Cloud Services

Haoran Yan, Yinfang Chen, Minghua Ma +10

The ever-increasing demand for generative artificial intelligence (GenAI) has motivated cloud-based GenAI services such as Azure OpenAI Service and Amazon Bedrock. Like any large-s…

cs.AI2025

ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Saurabh Jha, Rohan Arora, Yuji Watanabe +40

Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a…