papers

Publications (11)

cs.DC2016

Extreme Scale-out SuperMUC Phase 2 - lessons learned

Nicolay Hammer, Ferdinand Jamitzky, Helmut Satzger +36

In spring 2015, the Leibniz Supercomputing Centre (Leibniz-Rechenzentrum, LRZ), installed their new Peta-Scale System SuperMUC Phase2. Selected users were invited for a 28 day extr…

cs.DC2013

Asynchronous MPI for the Masses

Markus Wittmann, Georg Hager, Thomas Zeiser +1

We present a simple library which equips MPI implementations with truly asynchronous non-blocking point-to-point operations, and which is independent of the underlying communicatio…

cs.DC2015

Building a fault tolerant application using the GASPI communication layer

Faisal Shahzad, Moritz Kreutzer, Thomas Zeiser +4

It is commonly agreed that highly parallel software on Exascale computers will suffer from many more runtime failures due to the decreasing trend in the mean time to failures (MTTF…

cs.DC2007

RZBENCH: Performance evaluation of current HPC architectures using low-level and application benchmarks

Georg Hager, Holger Stengel, Thomas Zeiser +1

RZBENCH is a benchmark suite that was specifically developed to reflect the requirements of scientific supercomputer users at the University of Erlangen-Nuremberg (FAU). It compris…

cs.DC2011

Comparison of different Propagation Steps for the Lattice Boltzmann Method

Markus Wittmann, Thomas Zeiser, Georg Hager +1

Several possibilities exist to implement the propagation step of the lattice Boltzmann method. This paper describes common implementations which are compared according to the numbe…

cs.DC2017

CRAFT: A library for easier application-level Checkpoint/Restart and Automatic Fault Tolerance

Faisal Shahzad, Jonas Thies, Moritz Kreutzer +3

In order to efficiently use the future generations of supercomputers, fault tolerance and power consumption are two of the prime challenges anticipated by the High Performance Comp…