collaborators

5 papers

cs.DC2026

Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures

Shekhar Suman, Xiaoyu Chu, Alexandru Iosup

By 2025, there are zettabytes of data generated every year. The size and complexity of modern large-scale computing infrastructures like High-Performance Computing (HPC) systems co…

cs.PF2026

Leveraging LLMs for Structured Information Extraction and Analysis from Cloud Incident Reports (Work In Progress Paper)

Xiaoyu Chu, Shashikant Ilager, Yizhen Zang +2

Incident management is essential to maintain the reliability and availability of cloud computing services. Cloud vendors typically disclose incident reports to the public, summariz…

cs.DC2025

Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services

Sacheendra Talluri, Dante Niewenhuis, Xiaoyu Chu +4

Cloud services are critical to society. However, their reliability is poorly understood. Towards solving the problem, we propose a standard repository for cloud uptime data. We pop…

cs.PF2025

FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents

Sándor Battaglini-Fischer, Nishanthi Srinivasan, Bálint László Szarvas +2

Large Language Model (LLM) services such as ChatGPT, DALLE, and Cursor have quickly become essential for society, businesses, and individuals, empowering applications such as chatb…

cs.PF2025

An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models

Xiaoyu Chu, Sacheendra Talluri, Qingxian Lu +1

People and businesses increasingly rely on public LLM services, such as ChatGPT, DALLE, and Claude. Understanding their outages, and particularly measuring their failure-recovery p…