activity
20242026
collaborators

6 papers

cs.DC2026

M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis

Radu Nicolae, Dante Niewenhuis, Sacheendra Talluri +1

Datacenters are vital to our digital society, but consume a considerable fraction of global electricity and demand is projected to increase. To improve their sustainability and per…

cs.DC2026

OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters

Dante Niewenhuis, Sacheendra Talluri, Alexandru Iosup +1

The need to reduce datacenter carbon footprint is urgent. While many sustainability techniques have been proposed, they are often evaluated in isolation, using limited setups or an…

cs.PF2026

Leveraging LLMs for Structured Information Extraction and Analysis from Cloud Incident Reports (Work In Progress Paper)

Xiaoyu Chu, Shashikant Ilager, Yizhen Zang +2

Incident management is essential to maintain the reliability and availability of cloud computing services. Cloud vendors typically disclose incident reports to the public, summariz…

cs.DC2025

Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services

Sacheendra Talluri, Dante Niewenhuis, Xiaoyu Chu +4

Cloud services are critical to society. However, their reliability is poorly understood. Towards solving the problem, we propose a standard repository for cloud uptime data. We pop…

cs.PF2025

An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models

Xiaoyu Chu, Sacheendra Talluri, Qingxian Lu +1

People and businesses increasingly rely on public LLM services, such as ChatGPT, DALLE, and Claude. Understanding their outages, and particularly measuring their failure-recovery p…

cs.DC2024

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

Xiaoyu Chu, Daniel Hofstätter, Shashikant Ilager +6

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science,…