activity
20172021
most citedReport of the HPC Correctness Summit, Jan 25--26, 2017, Washington, DC

3 citations · 6 across the 4 of their papers we have counts for

collaborators
Showing cs.DCShow all

5 papers · 1 filter

cs.DC2021

Reinit++: Evaluating the Performance of Global-Restart Recovery Methods For MPI Fault Tolerance

Giorgis Georgakoudis, Luanzheng Guo, Ignacio Laguna

Scaling supercomputers comes with an increase in failure rates due to the increasing number of hardware components. In standard practice, applications are made resilient through ch…

cs.DC2021

MATCH: An MPI Fault Tolerance Benchmark Suite

Luanzheng Guo, Giorgis Georgakoudis, Konstantinos Parasyris +2

MPI has been ubiquitously deployed in flagship HPC systems aiming to accelerate distributed scientific applications running on tens of hundreds of processes and compute nodes. Main…

cs.DC2018

PARIS: Predicting Application Resilience Using Machine Learning

Luanzheng Guo, Dong Li, Ignacio Laguna

Extreme-scale scientific applications can be more vulnerable to soft errors (transient faults) as high-performance computing systems increase in scale. The common practice to evalu…

cs.DC2018

FlipTracker: Understanding Natural Error Resilience in HPC Applications

Luanzheng Guo, Dong Li, Ignacio Laguna +1

As high-performance computing systems scale in size and computational power, the danger of silent errors, i.e., errors that can bypass hardware detection mechanisms and impact appl…

cs.DC20173 cited

Report of the HPC Correctness Summit, Jan 25--26, 2017, Washington, DC

Ganesh Gopalakrishnan, Paul D. Hovland, Costin Iancu +6

Maintaining leadership in HPC requires the ability to support simulations at large scales and fidelity. In this study, we detail one of the most significant productivity challenges…