papers
Publications (11)
cs.DC2022
Implicit Actions and Non-blocking Failure Recovery with MPI
Aurelien Bouteiller, George Bosilca
cs.DC2008
Algorithmic Based Fault Tolerance Applied to High Performance Computing
George Bosilca, Remi Delmas, Jack Dongarra +1
cs.DC2021
Callback-based Completion Notification using MPI Continuations
Joseph Schuchart, Philipp Samfass, Christoph Niethammer +2
cs.DC2020
Task Bench: A Parameterized Benchmark for Evaluating Parallel Runtime Performance
Elliott Slaughter, Wei Wu, Yuankun Fu +12
cs.DC2021
Quo Vadis MPI RMA? Towards a More Efficient Use of MPI One-Sided Communication
Joseph Schuchart, Christoph Niethammer, José Gracia +1
cs.PF2023
Cache Optimization and Performance Modeling of Batched, Small, and Rectangular Matrix Multiplication on Intel, AMD, and Fujitsu Processors
Sameer Deshmukh, Rio Yokota, George Bosilca
math.NA2023
$O(N)$ distributed direct factorization of structured dense matrices using runtime systems
Sameer Deshmukh, Qinxiang Ma, Rio Yokota +1
cs.DC2021
SuperNeurons: FFT-based Gradient Sparsification in the Distributed Training of Deep Neural Networks
Linnan Wang, Wei Wu, Junyu Zhang +4
cs.DC2017
Efficient Communications in Training Large Scale Neural Networks
Linnan Wang, Wei Wu, George Bosilca +2
stat.CO2024
Boosting Earth System Model Outputs And Saving PetaBytes in their Storage Using Exascale Climate Emulators
Sameh Abdulah, Allison H. Baker, George Bosilca +9
cs.DC2014
Taking advantage of hybrid systems for sparse direct solvers via task-based runtimes
Xavier Lacoste, Mathieu Faverge, George Bosilca +2