3 citations · 3 across the 1 of their papers we have counts for
4 papers · 1 filter
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
Mikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek +3
In the Fully Sharded Data Parallel (FSDP) training pipeline, collective operations can be interleaved to maximize the communication/computation overlap. In this scenario, outstandi…
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
Luigi Fusco, Mikhail Khalilov, Marcin Chrapek +3
Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads…
Software Resource Disaggregation for HPC with Serverless Computing
Marcin Copik, Marcin Chrapek, Larissa Schmid +2
Aggregated HPC resources have rigid allocation systems and programming models which struggle to adapt to diverse and changing workloads. Consequently, HPC systems fail to efficient…
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
Siyuan Shen, Langwen Huang, Marcin Chrapek +5
The shift towards high-bandwidth networks driven by AI workloads in data centers and HPC clusters has unintentionally aggravated network latency, adversely affecting the performanc…