40 citations · 95 across the 12 of their papers we have counts for
7 papers · 1 filter
Anomaly Detection using Autoencoders in High Performance Computing Systems
Andrea Borghesi, Andrea Bartolini, Michele Lombardi +2
Anomaly detection in supercomputers is a very difficult problem due to the big scale of the systems and the high number of components. The current state of the art for automated an…
Robust identification of thermal models for in-production High-Performance-Computing clusters with machine learning-based data selection
Federico Pittino, Roberto Diversi, Luca Benini +1
Power and thermal management are critical components of High-Performance-Computing (HPC) systems, due to their high power density and large total power consumption. The assessment…
Online Fault Classification in HPC Systems through Machine Learning
Alessio Netti, Zeynep Kiziltan, Ozalp Babaoglu +3
As High-Performance Computing (HPC) systems strive towards the exascale goal, studies suggest that they will experience excessive failure rates. For this reason, detecting and clas…
FINJ: A Fault Injection Tool for HPC Systems
Alessio Netti, Zeynep Kiziltan, Ozalp Babaoglu +3
We present FINJ, a high-level fault injection tool for High-Performance Computing (HPC) systems, with a focus on the management of complex experiments. FINJ provides support for cu…
Pricing Schemes for Energy-Efficient HPC Systems: Design and Exploration
Andrea Borghesi, Andrea Bartolini, Michela Milano +1
Energy efficiency is of paramount importance for the sustainability of HPC systems. Energy consumption limits the peak performance of supercomputers and accounts for a large share…
COUNTDOWN: a Run-time Library for Performance-Neutral Energy Saving in MPI Applications
Daniele Cesarini, Andrea Bartolini, Pietro Bonfà +2
Power and energy consumption is becoming key challenges to deploy the first exascale supercomputer successfully. Large-scale HPC applications waste a significant amount of power in…