activity
20162023
most citedSpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

26 citations · 39 across the 10 of their papers we have counts for

collaborators
Showing cs.DCShow all

11 papers · 1 filter

cs.DC20245 cited

Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip

Luigi Fusco, Mikhail Khalilov, Marcin Chrapek +3

Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads…

cs.DC2024

LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming

Siyuan Shen, Langwen Huang, Marcin Chrapek +5

The shift towards high-bandwidth networks driven by AI workloads in data centers and HPC clusters has unintentionally aggravated network latency, adversely affecting the performanc…

cs.DC20246 cited

Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication

Lukas Gianinazzi, Alexandros Nikolaos Ziogas, Langwen Huang +9

We propose a novel approach to iterated sparse matrix dense matrix multiplication, a fundamental computational kernel in scientific computing and graph neural network training. In…

cs.DC20242 cited

XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing

Torsten Hoefler, Marcin Copik, Pete Beckman +8

HPC and Cloud have evolved independently, specializing their innovations into performance or productivity. Acceleration as a Service (XaaS) is a recipe to empower both fields with…

cs.DC20243 cited

Swing: Short-cutting Rings for Higher Bandwidth Allreduce

Daniele De Sensi, Tommaso Bonato, David Saam +1

The allreduce collective operation accounts for a significant fraction of the runtime of workloads running on distributed systems. One factor determining its performance is the dis…

cs.DC2023

VENOM: A Vectorized N:M Format for Unleashing the Power of Sparse Tensor Cores

Roberto L. Castro, Andrei Ivanov, Diego Andrade +3

The increasing success and scaling of Deep Learning models demands higher computational efficiency and power. Sparsification can lead to both smaller models as well as higher compu…