5 papers
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
Shekhar Suman, Xiaoyu Chu, Alexandru Iosup
By 2025, there are zettabytes of data generated every year. The size and complexity of modern large-scale computing infrastructures like High-Performance Computing (HPC) systems co…
Leveraging LLMs for Structured Information Extraction and Analysis from Cloud Incident Reports (Work In Progress Paper)
Xiaoyu Chu, Shashikant Ilager, Yizhen Zang +2
Incident management is essential to maintain the reliability and availability of cloud computing services. Cloud vendors typically disclose incident reports to the public, summariz…
Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services
Sacheendra Talluri, Dante Niewenhuis, Xiaoyu Chu +4
Cloud services are critical to society. However, their reliability is poorly understood. Towards solving the problem, we propose a standard repository for cloud uptime data. We pop…
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
Sándor Battaglini-Fischer, Nishanthi Srinivasan, Bálint László Szarvas +2
Large Language Model (LLM) services such as ChatGPT, DALLE, and Cursor have quickly become essential for society, businesses, and individuals, empowering applications such as chatb…
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
Xiaoyu Chu, Sacheendra Talluri, Qingxian Lu +1
People and businesses increasingly rely on public LLM services, such as ChatGPT, DALLE, and Claude. Understanding their outages, and particularly measuring their failure-recovery p…