activity
20182021
collaborators
Showing cs.DCShow all

7 papers · 1 filter

cs.DC202412 cited

On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems

Eishi Arima, Isaías A. Comprés, Martin Schulz

Recent High-Performance Computing (HPC) systems are facing important challenges, such as massive power consumption, while at the same time significantly under-utilized system resou…

cs.DC2021

Operational Data Analytics in Practice: Experiences from Design to Deployment in Production HPC Environments

Alessio Netti, Michael Ott, Carla Guillen +2

As HPC systems grow in complexity, efficient and manageable operation is increasingly critical. Many centers are thus starting to explore the use of Operational Data Analytics (ODA…

cs.DC2020

Resiliency in Numerical Algorithm Design for Extreme Scale Simulations

Emmanuel Agullo, Mirco Altenbernd, Hartwig Anzt +33

This work is based on the seminar titled ``Resiliency in Numerical Algorithm Design for Extreme Scale Simulations'' held March 1-6, 2020 at Schloss Dagstuhl, that was attended by a…

cs.DC2020

Correlation-wise Smoothing: Lightweight Knowledge Extraction for HPC Monitoring Data

Alessio Netti, Daniele Tafani, Michael Ott +1

Modern High-Performance Computing (HPC) and data center operators rely more and more on data analytics techniques to improve the efficiency and reliability of their operations. The…

cs.DC2019

DCDB Wintermute: Enabling Online and Holistic Operational Data Analytics on HPC Systems

Alessio Netti, Micha Mueller, Carla Guillen +4

As we approach the exascale era, the size and complexity of HPC systems continues to increase, raising concerns about their manageability and sustainability. For this reason, more…

cs.DC2019

From Facility to Application Sensor Data: Modular, Continuous and Holistic Monitoring with DCDB

Alessio Netti, Micha Mueller, Axel Auweter +4

Today's HPC installations are highly-complex systems, and their complexity will only increase as we move to exascale and beyond. At each layer, from facilities to systems, from run…