27 citations · 46 across the 11 of their papers we have counts for
13 papers · 1 filter
MOARD: Modeling Application Resilience to Transient Faults on Data Objects
Luanzheng Guo, Dong Li
Understanding application resilience (or error tolerance) in the presence of hardware transient faults on data objects is critical to ensure computing integrity and enable efficien…
MATCH: An MPI Fault Tolerance Benchmark Suite
Luanzheng Guo, Giorgis Georgakoudis, Konstantinos Parasyris +2
MPI has been ubiquitously deployed in flagship HPC systems aiming to accelerate distributed scientific applications running on tens of hundreds of processes and compute nodes. Main…
Demystifying the Performance of HPC Scientific Applications on NVM-based Memory Systems
Ivy Peng, Kai Wu, Jie Ren +2
The emergence of high-density byte-addressable non-volatile memory (NVM) is promising to accelerate data- and compute-intensive applications. Current NVM technologies have lower pe…
UMap: Enabling Application-driven Optimizations for Page Management
Ivy B. Peng, Marty McFadden, Eric Green +5
Leadership supercomputers feature a diversity of storage, from node-local persistent memory and NVMe SSDs to network-interconnected flash memory and HDD. Memory mapping files on di…
Architecture-Aware, High Performance Transaction for Persistent Memory
Kai Wu, Jie Ren, Dong Li
Byte-addressable non-volatile main memory (NVM) demands transactional mechanisms to access and manipulate data on NVM atomically. Those transaction mechanisms often employ a loggin…
PARIS: Predicting Application Resilience Using Machine Learning
Luanzheng Guo, Dong Li, Ignacio Laguna
Extreme-scale scientific applications can be more vulnerable to soft errors (transient faults) as high-performance computing systems increase in scale. The common practice to evalu…