5 citations · 5 across the 2 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2024
Checkpoint and Restart: An Energy Consumption Characterization in Clusters
Marina Moran, Javier Balladini, Dolores Rexachs +1
The fault tolerance method currently used in High Performance Computing (HPC) is the rollback-recovery method by using checkpoints. This, like any other fault tolerance method, add…
cs.DC2020★ 5 cited
Soft Errors Detection and Automatic Recovery based on Replication combined with different Levels of Checkpointing
Diego Montezanti, Enzo Rucci, Armando De Giusti +3
Handling faults is a growing concern in HPC. In future exascale systems, it is projected that silent undetected errors will occur several times a day, increasing the occurrence of…