3 citations · 6 across the 4 of their papers we have counts for
5 papers · 1 filter
Reinit++: Evaluating the Performance of Global-Restart Recovery Methods For MPI Fault Tolerance
Giorgis Georgakoudis, Luanzheng Guo, Ignacio Laguna
Scaling supercomputers comes with an increase in failure rates due to the increasing number of hardware components. In standard practice, applications are made resilient through ch…
MATCH: An MPI Fault Tolerance Benchmark Suite
Luanzheng Guo, Giorgis Georgakoudis, Konstantinos Parasyris +2
MPI has been ubiquitously deployed in flagship HPC systems aiming to accelerate distributed scientific applications running on tens of hundreds of processes and compute nodes. Main…
PARIS: Predicting Application Resilience Using Machine Learning
Luanzheng Guo, Dong Li, Ignacio Laguna
Extreme-scale scientific applications can be more vulnerable to soft errors (transient faults) as high-performance computing systems increase in scale. The common practice to evalu…
FlipTracker: Understanding Natural Error Resilience in HPC Applications
Luanzheng Guo, Dong Li, Ignacio Laguna +1
As high-performance computing systems scale in size and computational power, the danger of silent errors, i.e., errors that can bypass hardware detection mechanisms and impact appl…
Report of the HPC Correctness Summit, Jan 25--26, 2017, Washington, DC
Ganesh Gopalakrishnan, Paul D. Hovland, Costin Iancu +6
Maintaining leadership in HPC requires the ability to support simulations at large scales and fidelity. In this study, we detail one of the most significant productivity challenges…