Publications (11)
Extreme Scale-out SuperMUC Phase 2 - lessons learned
Nicolay Hammer, Ferdinand Jamitzky, Helmut Satzger +36
In spring 2015, the Leibniz Supercomputing Centre (Leibniz-Rechenzentrum, LRZ), installed their new Peta-Scale System SuperMUC Phase2. Selected users were invited for a 28 day extr…
Asynchronous MPI for the Masses
Markus Wittmann, Georg Hager, Thomas Zeiser +1
We present a simple library which equips MPI implementations with truly asynchronous non-blocking point-to-point operations, and which is independent of the underlying communicatio…
Building a fault tolerant application using the GASPI communication layer
Faisal Shahzad, Moritz Kreutzer, Thomas Zeiser +4
It is commonly agreed that highly parallel software on Exascale computers will suffer from many more runtime failures due to the decreasing trend in the mean time to failures (MTTF…
RZBENCH: Performance evaluation of current HPC architectures using low-level and application benchmarks
Georg Hager, Holger Stengel, Thomas Zeiser +1
RZBENCH is a benchmark suite that was specifically developed to reflect the requirements of scientific supercomputer users at the University of Erlangen-Nuremberg (FAU). It compris…
Comparison of different Propagation Steps for the Lattice Boltzmann Method
Markus Wittmann, Thomas Zeiser, Georg Hager +1
Several possibilities exist to implement the propagation step of the lattice Boltzmann method. This paper describes common implementations which are compared according to the numbe…
CRAFT: A library for easier application-level Checkpoint/Restart and Automatic Fault Tolerance
Faisal Shahzad, Jonas Thies, Moritz Kreutzer +3
In order to efficiently use the future generations of supercomputers, fault tolerance and power consumption are two of the prime challenges anticipated by the High Performance Comp…