activity
20152020
most citedA Machine Learning Approach to Online Fault Classification in HPC Systems

28 citations · 28 across the 1 of their papers we have counts for

collaborators
Showing cs.DCShow all

6 papers · 1 filter

cs.DC202028 cited

A Machine Learning Approach to Online Fault Classification in HPC Systems

Alessio Netti, Zeynep Kiziltan, Ozalp Babaoglu +3

As High-Performance Computing (HPC) systems strive towards the exascale goal, failure rates both at the hardware and software levels will increase significantly. Thus, detecting an…

cs.DC2018

Online Fault Classification in HPC Systems through Machine Learning

Alessio Netti, Zeynep Kiziltan, Ozalp Babaoglu +3

As High-Performance Computing (HPC) systems strive towards the exascale goal, studies suggest that they will experience excessive failure rates. For this reason, detecting and clas…

cs.DC2018

FINJ: A Fault Injection Tool for HPC Systems

Alessio Netti, Zeynep Kiziltan, Ozalp Babaoglu +3

We present FINJ, a high-level fault injection tool for High-Performance Computing (HPC) systems, with a focus on the management of complex experiments. FINJ provides support for cu…

cs.DC2016

Towards Operator-less Data Centers Through Data-Driven, Predictive, Proactive Autonomics

Alina Sîrbu, Ozalp Babaoglu

Continued reliance on human operators for managing data centers is a major impediment for them from ever reaching extreme dimensions. Large computer systems in general, and data ce…

cs.DC2016

Predicting System-level Power for a Hybrid Supercomputer

Alina Sîrbu, Ozalp Babaoglu

For current High Performance Computing systems to scale towards the holy grail of ExaFLOP performance, their power consumption has to be reduced by at least one order of magnitude.…

cs.DC2015

Towards Data-Driven Autonomics in Data Centers

Alina Sîrbu, Ozalp Babaoglu

Continued reliance on human operators for managing data centers is a major impediment for them from ever reaching extreme dimensions. Large computer systems in general, and data ce…