26 citations · 39 across the 10 of their papers we have counts for
11 papers · 1 filter
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
Luigi Fusco, Mikhail Khalilov, Marcin Chrapek +3
Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads…
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
Siyuan Shen, Langwen Huang, Marcin Chrapek +5
The shift towards high-bandwidth networks driven by AI workloads in data centers and HPC clusters has unintentionally aggravated network latency, adversely affecting the performanc…
Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication
Lukas Gianinazzi, Alexandros Nikolaos Ziogas, Langwen Huang +9
We propose a novel approach to iterated sparse matrix dense matrix multiplication, a fundamental computational kernel in scientific computing and graph neural network training. In…
XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing
Torsten Hoefler, Marcin Copik, Pete Beckman +8
HPC and Cloud have evolved independently, specializing their innovations into performance or productivity. Acceleration as a Service (XaaS) is a recipe to empower both fields with…
Swing: Short-cutting Rings for Higher Bandwidth Allreduce
Daniele De Sensi, Tommaso Bonato, David Saam +1
The allreduce collective operation accounts for a significant fraction of the runtime of workloads running on distributed systems. One factor determining its performance is the dis…
VENOM: A Vectorized N:M Format for Unleashing the Power of Sparse Tensor Cores
Roberto L. Castro, Andrei Ivanov, Diego Andrade +3
The increasing success and scaling of Deep Learning models demands higher computational efficiency and power. Sparsification can lead to both smaller models as well as higher compu…