5 papers · 1 filter
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
Baodi Shan, Mauricio Araya-Polo, Barbara Chapman
Distributed GPU applications increasingly rely on kernel-level, cross-node coordination to reduce launch overheads and improve compute-communication overlap, but such support is la…
DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP
Baodi Shan, Mauricio Araya-Polo, Barbara Chapman
As core counts and heterogeneity rise in HPC, traditional hybrid programming models face challenges in managing distributed GPU memory and ensuring portability. This paper presents…
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
Baodi Shan, Mauricio Araya-Polo, Barbara Chapman
MPI+X has been the de facto standard for distributed memory parallel programming. It is widely used primarily as an explicit two-sided communication model, which often leads to com…
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
Baodi Shan, Mauricio Araya-Polo
Accelerated computing is widely used in high-performance computing. Therefore, it is crucial to experiment and discover how to better utilize GPUGPUs latest generations on relevant…
LCFI: A Fault Injection Tool for Studying Lossy Compression Error Propagation in HPC Programs
Baodi Shan, Aabid Shamji, Jiannan Tian +2
Error-bounded lossy compression is becoming more and more important to today's extreme-scale HPC applications because of the ever-increasing volume of data generated because it has…