papers

Publications (260)

cs.NI2025

Ultra Ethernet's Design Principles and Architectural Innovations

Torsten Hoefler, Karen Schramm, Eric Spada +13

The recently released Ultra Ethernet (UE) 1.0 specification defines a transformative High-Performance Ethernet standard for future Artificial Intelligence (AI) and High-Performance…

cs.DB2026

Benchmarking Filtered Approximate Nearest Neighbor Search Algorithms on Transformer-based Embedding Vectors

Patrick Iff, Paul Bruegger, Marcin Chrapek +3

Advances in embedding models for text, image, audio, and video drive progress across multiple domains, including retrieval-augmented generation, recommendation systems, and others.…

cs.DC2024

Swing: Short-cutting Rings for Higher Bandwidth Allreduce

Daniele De Sensi, Tommaso Bonato, David Saam +1

The allreduce collective operation accounts for a significant fraction of the runtime of workloads running on distributed systems. One factor determining its performance is the dis…

cs.DS2026

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding

Jiale Chen, Torsten Hoefler, Dan Alistarh

The paper introduces GPTQ-2D, an algorithm that rounds a real matrix to integers under a two-sided quadratic metric in cubic time by processing entries anti-diagonal by anti-diagon…

#matrix rounding#adaptive rounding#quadratic metric#algorithmic complexity
cs.LG2023

Cached Operator Reordering: A Unified View for Fast GNN Training

Julia Bazinska, Andrei Ivanov, Tal Ben-Nun +4

Graph Neural Networks (GNNs) are a powerful tool for handling structured graph data and addressing tasks such as node classification, graph classification, and clustering. However,…

cs.PL2020

Stateful Dataflow Multigraphs: A Data-Centric Model for Performance Portability on Heterogeneous Architectures

Tal Ben-Nun, Johannes de Fine Licht, Alexandros Nikolaos Ziogas +2

The ubiquity of accelerators in high-performance computing has driven programming complexity beyond the skill-set of the average domain scientist. To maintain performance portabili…

cs.DC2026

Demystifying NVSHMEM: A System-Level Analysis on Symmetric Memory and Device-Initiated Operations in GPU Communication

Yijun Ma, Siyuan Shen, Tiancheng Chen +6

NVSHMEM is NVIDIA's OpenSHMEM-based PGAS communication library for GPU clusters, enabling GPU-initiated, one-sided communication through symmetric memory. Despite its growing adopt…

cs.DC2025

Ab-initio Quantum Transport with the GW Approximation, 42,240 Atoms, and Sustained Exascale Performance

Nicolas Vetsch, Alexander Maeder, Vincent Maillou +7

Designing nanoscale electronic devices such as the currently manufactured nanoribbon field-effect transistors (NRFETs) requires advanced modeling tools capturing all relevant quant…

cs.AR2025

RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures

Patrick Iff, Benigna Bruggmann, Blaise Morel +3

Chiplet architectures are on the rise as they promise to overcome the scaling challenges of monolithic chips. A key component of such architectures is an efficient inter-chiplet in…

cs.LG2021

Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks

Torsten Hoefler, Dan Alistarh, Tal Ben-Nun +2

The growing energy and performance costs of deep learning have driven the community to reduce the size of neural networks by selectively pruning components. Similarly to their biol…

physics.comp-ph2019

A scalable weakly-synchronous algorithm for solving partial differential equations

Konduri Aditya, Tobias Gysi, Grzegorz Kwasniewski +3

Synchronization overheads pose a major challenge as applications advance towards extreme scales. In current large-scale algorithms, synchronization as well as data communication de…

cs.DC2021

Flare: Flexible In-Network Allreduce

Daniele De Sensi, Salvatore Di Girolamo, Saleh Ashkboos +2

The allreduce operation is one of the most commonly used communication routines in distributed applications. To improve its bandwidth and to reduce network traffic, this operation…

cs.DC2024

Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI

Mikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek +3

In the Fully Sharded Data Parallel (FSDP) training pipeline, collective operations can be interleaved to maximize the communication/computation overlap. In this scenario, outstandi…

cs.DC2020

Extracting Clean Performance Models from Tainted Programs

Marcin Copik, Alexandru Calotoiu, Tobias Grosser +3

Performance models are well-known instruments to understand the scaling behavior of parallel applications. They express how performance changes as key execution parameters, such as…

cs.DB2023

Demystifying Graph Databases: Analysis and Taxonomy of Data Organization, System Designs, and Graph Queries

Maciej Besta, Robert Gerstenberger, Emanuel Peter +5

Graph processing has become an important part of multiple areas of computer science, such as machine learning, computational sciences, medical applications, social network analysis…

cs.DS2021

Slim Graph: Practical Lossy Graph Compression for Approximate Graph Processing, Storage, and Analytics

Maciej Besta, Simon Weber, Lukas Gianinazzi +4

We propose Slim Graph: the first programming model and framework for practical lossy graph compression that facilitates high-performance approximate graph processing, storage, and…

cs.CE2019

A Data-Centric Approach to Extreme-Scale Ab initio Dissipative Quantum Transport Simulations

Alexandros Nikolaos Ziogas, Tal Ben-Nun, Guillermo Indalecio Fernández +3

The computational efficiency of a state of the art ab initio quantum transport (QT) solver, capable of revealing the coupled electro-thermal properties of atomically-resolved nano-…

cs.DB2026

GraphSeek: Next-Generation Graph Analytics with LLMs

Maciej Besta, Łukasz Jarmocik, Orest Hrycyna +7

Graphs are foundational across domains but remain hard to use without deep expertise. LLMs promise accessible natural language (NL) graph analytics, yet they fail to process indust…

cs.LG2024

QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci +6

We introduce QuaRot, a new Quantization scheme based on Rotations, which is able to quantize LLMs end-to-end, including all weights, activations, and KV cache in 4 bits. QuaRot rot…

cs.LG2025

Demystifying Higher-Order Graph Neural Networks

Maciej Besta, Florian Scheidl, Lukas Gianinazzi +4

Higher-order graph neural networks (HOGNNs) and the related architectures from Topological Deep Learning are an important class of GNN models that harness polyadic relations betwee…

cs.DC2023

Efficient Quantized Sparse Matrix Operations on Tensor Cores

Shigang Li, Kazuki Osawa, Torsten Hoefler

The exponentially growing model size drives the continued success of deep learning, but it brings prohibitive computation and memory cost. From the algorithm perspective, model spa…

cs.NI2021

PsPIN: A high-performance low-power architecture for flexible in-network compute

Salvatore Di Girolamo, Andreas Kurth, Alexandru Calotoiu +5

The capacity of offloading data and control tasks to the network is becoming increasingly important, especially if we consider the faster growth of network speed when compared to C…

cs.CR2024

Fortify Your Foundations: Practical Privacy and Security for Foundation Model Deployments In The Cloud

Marcin Chrapek, Anjo Vahldiek-Oberwagner, Marcin Spoczynski +3

Foundation Models (FMs) display exceptional performance in tasks such as natural language processing and are being applied across a growing range of disciplines. Although typically…

cs.DS2021

Parallel Algorithms for Finding Large Cliques in Sparse Graphs

Lukas Gianinazzi, Maciej Besta, Yannick Schaffner +1

We present a parallel k-clique listing algorithm with improved work bounds (for the same depth) in sparse graphs with low degeneracy or arboricity. We achieve this by introducing a…

cs.AR2023

Sparse Stream Semantic Registers: A Lightweight ISA Extension Accelerating General Sparse Linear Algebra

Paul Scheffler, Florian Zaruba, Fabian Schuiki +2

Sparse linear algebra is crucial in many application domains, but challenging to handle efficiently in both software and hardware, with one- and two-sided operand sparsity handled…

cs.DC2024

XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing

Torsten Hoefler, Marcin Copik, Pete Beckman +8

HPC and Cloud have evolved independently, specializing their innovations into performance or productivity. Acceleration as a Service (XaaS) is a recipe to empower both fields with…

cs.LG2024

All models are wrong, some are useful: Model Selection with Limited Labels

Patrik Okanovic, Andreas Kirsch, Jannes Kasper +3

We introduce MODEL SELECTOR, a framework for label-efficient selection of pretrained classifiers. Given a pool of unlabeled target data, MODEL SELECTOR samples a small subset of hi…

cs.DC2021

Practice of Streaming Processing of Dynamic Graphs: Concepts, Models, and Systems

Maciej Besta, Marc Fischer, Vasiliki Kalavri +2

Graph processing has become an important part of various areas of computing, including machine learning, medical applications, social network analysis, computational sciences, and…

cs.DC2022

Deinsum: Practically I/O Optimal Multilinear Algebra

Alexandros Nikolaos Ziogas, Grzegorz Kwasniewski, Tal Ben-Nun +2

Multilinear algebra kernel performance on modern massively-parallel systems is determined mainly by data movement. However, deriving data movement-optimal distributed schedules for…

cs.DC2019

Streaming Message Interface: High-Performance Distributed Memory Programming on Reconfigurable Hardware

Tiziano De Matteis, Johannes de Fine Licht, Jakub Beránek +1

Distributed memory programming is the established paradigm used in high-performance computing (HPC) systems, requiring explicit communication between nodes and devices. When FPGAs…

cs.AR2023

HexaMesh: Scaling to Hundreds of Chiplets with an Optimized Chiplet Arrangement

Patrick Iff, Maciej Besta, Matheus Cavalcante +3

2.5D integration is an important technique to tackle the growing cost of manufacturing chips in advanced technology nodes. This poses the challenge of providing high-performance in…

cs.PF2025

EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC

Siyuan Shen, Mikhail Khalilov, Lukas Gianinazzi +6

Resource disaggregation is a promising technique for improving the efficiency of large-scale computing systems. However, this comes at the cost of increased memory access latency d…

cs.CE2024

DiffDA: a Diffusion Model for Weather-scale Data Assimilation

Langwen Huang, Lukas Gianinazzi, Yuejiang Yu +2

The generation of initial conditions via accurate data assimilation is crucial for weather forecasting and climate modeling. We propose DiffDA as a denoising diffusion model capabl…

cs.LG2026

The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm

Jiale Chen, Yalda Shabanzadeh, Elvir Crnčević +2

Quantizing the weights of large language models (LLMs) from 16-bit to lower bitwidth is the de facto approach to deploy massive transformers onto more affordable accelerators. Whil…

cs.LG2025

DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing

Afif Boudaoud, Alexandru Calotoiu, Marcin Copik +1

Automatic differentiation (AD) is a set of techniques that systematically applies the chain rule to compute the gradients of functions without requiring human intervention. Althoug…

cs.DC2025

Core Hours and Carbon Credits: Incentivizing Sustainability in HPC

Alok Kamatar, Maxime Gonthier, Valerie Hayot-Sasson +6

Realizing a shared responsibility between providers and consumers is critical to manage the sustainability of HPC. However, while cost may motivate efficiency improvements by infra…

cs.CE2026

Error bounded compression for weather and climate applications

Langwen Huang, Luigi Fusco, Florian Scheidl +4

As the resolution of weather and climate simulations increases, the amount of data produced is growing rapidly from hundreds of terabytes to tens of petabytes. The huge size become…

cs.AR2025

Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal

Wenqi Jiang, Hang Hu, Torsten Hoefler +1

Vector search systems are indispensable in large language model (LLM) serving, search engines, and recommender systems, where minimizing online search latency is essential. Among v…

cs.PL2026

MLIR-Forge: A Modular Framework for Language Smiths

Berke Ates, Philipp Schaad, Timo Schneider +2

Optimizing compilers are essential for the efficient and correct execution of software across various scientific fields. Domain-specific languages (DSL) typically use higher level…

cs.LG2018

The Convergence of Sparsified Gradient Methods

Dan Alistarh, Torsten Hoefler, Mikael Johansson +3

Distributed training of massive machine learning models, in particular deep neural networks, via Stochastic Gradient Descent (SGD) is becoming commonplace. Several families of comm…

cs.NI2024

A High-Performance Design, Implementation, Deployment, and Evaluation of The Slim Fly Network

Nils Blach, Maciej Besta, Daniele De Sensi +10

Novel low-diameter network topologies such as Slim Fly (SF) offer significant cost and power advantages over the established Fat Tree, Clos, or Dragonfly. To spearhead the adoption…

cs.DC2020

Active Access: A Mechanism for High-Performance Distributed Data-Centric Computations

Maciej Besta, Torsten Hoefler

Remote memory access (RMA) is an emerging high-performance programming model that uses RDMA hardware directly. Yet, accessing remote memories cannot invoke activities at the target…

cs.DC2026

SpaDA: A Spatial Dataflow Architecture Programming Language

Lukas Gianinazzi, Tal Ben-Nun, Torsten Hoefler

Spatial dataflow architectures like the Cerebras Wafer-Scale Engine deliver exceptional performance in AI and scientific computing by distributing scratchpad memory across hundreds…

cs.AI2025

Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge

Drago Plecko, Patrik Okanovic, Shreyas Havaldar +2

Artificial intelligence (AI) systems hold great promise for advancing various scientific disciplines, and are increasingly used in real-world applications. Despite their remarkable…

cs.LG2024

MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Elias Frantar, Roberto L. Castro, Jiale Chen +2

As inference on Large Language Models (LLMs) emerges as an important workload in machine learning applications, weight quantization has become a standard technique for efficient GP…

cs.DC2024

LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming

Siyuan Shen, Langwen Huang, Marcin Chrapek +5

The shift towards high-bandwidth networks driven by AI workloads in data centers and HPC clusters has unintentionally aggravated network latency, adversely affecting the performanc…

cs.DC2025

Evolving HPC services to enable ML workloads on HPE Cray EX

Stefano Schuppli, Fawzi Mohamed, Henrique Mendonça +10

The Alps Research Infrastructure leverages GH200 technology at scale, featuring 10,752 GPUs. Accessing Alps provides a significant computational advantage for researchers in Artifi…

cs.LG2022

A Data-Centric Optimization Framework for Machine Learning

Oliver Rausch, Tal Ben-Nun, Nikoli Dryden +3

Rapid progress in deep learning is leading to a diverse set of quickly changing models, with a dramatically growing demand for compute. However, as frameworks specialize performanc…

cs.NI2024

FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network

Timo Schneider, Pengcheng Xu, Torsten Hoefler

In the era of post-Moore computing, network offload emerges as a solution to two challenges: the imperative for low-latency communication and the push towards hardware specialisati…

cs.DC2021

Communication Lower Bounds of Bilinear Algorithms for Symmetric Tensor Contractions

Edgar Solomonik, James Demmel, Torsten Hoefler

We introduce a new theoretical framework for deriving lower bounds on data movement in bilinear algorithms. Bilinear algorithms are a general representation of fast algorithms for…

cs.CE2019

Optimizing the Data Movement in Quantum Transport Simulations via Data-Centric Parallel Programming

Alexandros Nikolaos Ziogas, Tal Ben-Nun, Guillermo Indalecio Fernández +3

Designing efficient cooling systems for integrated circuits (ICs) relies on a deep understanding of the electro-thermal properties of transistors. To shed light on this issue in cu…

cs.PF2020

ScalAna: Automating Scaling Loss Detection with Graph Analysis

Yuyang Jin, Haojie Wang, Teng Yu +4

Scaling a parallel program to modern supercomputers is challenging due to inter-process communication, Amdahl's law, and resource contention. Performance analysis tools for finding…

cs.PF2025

Denoising Application Performance Models with Noise-Resilient Priors

Gustavo de Morais, Alexander Geiß, Alexandru Calotoiu +5

As parallel codes are scaled to larger computing systems, performance models play a crucial role in identifying potential bottlenecks. However, constructing these models analytical…

cs.DC2024

Near-Optimal Wafer-Scale Reduce

Piotr Luczynski, Lukas Gianinazzi, Patrick Iff +3

Efficient Reduce and AllReduce communication collectives are a critical cornerstone of high-performance computing (HPC) applications. We present the first systematic investigation…

cs.DC2022

Temporal Vectorization: A Compiler Approach to Automatic Multi-Pumping

Carl-Johannes Johnsen, Tiziano De Matteis, Tal Ben-Nun +2

The multi-pumping resource sharing technique can overcome the limitations commonly found in single-clocked FPGA designs by allowing hardware components to operate at a higher clock…

cs.LG2022

Spatial Mixture-of-Experts

Nikoli Dryden, Torsten Hoefler

Many data have an underlying dependence on spatial location; it may be weather on the Earth, a simulation on a mesh, or a registered image. Yet this feature is rarely taken advanta…

cs.DC2025

Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators

Aofeng Shen, Chi Zhang, Yakup Budanaz +4

Tile-based many-Processing Element (PE) accelerators can achieve competitive performance on General Matrix Multiplication (GEMM), but they are extremely hard to program, as their o…

cs.LG2026

When Data Is Scarce: Scaling Sparse Language Models with Repeated Training

Boqian Wu, Qiao Xiao, Patrik Okanovic +6

Scaling laws for dense LLMs under infinite data are well explored, but how sparsity interacts with limited data is not. In this work, we study sparse training in data-constrained r…

cs.SI2022

Motif Prediction with Graph Neural Networks

Maciej Besta, Raphael Grob, Cesare Miglioli +8

Link prediction is one of the central problems in graph mining. However, recent studies highlight the importance of higher-order network analysis, where complex structures called m…

cs.IR2024

Hardware Acceleration for Knowledge Graph Processing: Challenges & Recent Developments

Maciej Besta, Robert Gerstenberger, Patrick Iff +9

Knowledge graphs (KGs) have achieved significant attention in recent years, particularly in the area of the Semantic Web as well as gaining popularity in other application domains…

cs.DC2024

Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip

Luigi Fusco, Mikhail Khalilov, Marcin Chrapek +3

Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads…

cs.NI2026

REPS: Recycled Entropy Packet Spraying for Adaptive Load Balancing and Failure Mitigation

Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini +7

Next-generation datacenters require highly efficient network load balancing to manage the growing scale of artificial intelligence (AI) training and general datacenter traffic. How…

cs.NI2024

OSMOSIS: Enabling Multi-Tenancy in Datacenter SmartNICs

Mikhail Khalilov, Marcin Chrapek, Siyuan Shen +7

Multi-tenancy is essential for unleashing SmartNIC's potential in datacenters. Our systematic analysis in this work shows that existing on-path SmartNICs have resource multiplexing…

cs.DC2021

Flexible Communication Avoiding Matrix Multiplication on FPGA with High-Level Synthesis

Johannes de Fine Licht, Grzegorz Kwasniewski, Torsten Hoefler

Data movement is the dominating factor affecting performance and energy in modern computing systems. Consequently, many algorithms have been developed to minimize the number of I/O…

cs.PF2020

A Fast Analytical Model of Fully Associative Caches

Tobias Gysi, Tobias Grosser, Laurin Brandner +1

While the cost of computation is an easy to understand local property, the cost of data movement on cached architectures depends on global state, does not compose, and is hard to p…

cs.AR2020

Stream Semantic Registers: A Lightweight RISC-V ISA Extension Achieving Full Compute Utilization in Single-Issue Cores

Fabian Schuiki, Florian Zaruba, Torsten Hoefler +1

Single-issue processor cores are very energy efficient but suffer from the von Neumann bottleneck, in that they must explicitly fetch and issue the loads/storse necessary to feed t…

cs.DC2021

StencilFlow: Mapping Large Stencil Programs to Distributed Spatial Computing Systems

Johannes de Fine Licht, Andreas Kuster, Tiziano De Matteis +3

Spatial computing devices have been shown to significantly accelerate stencil computations, but have so far relied on unrolling the iterative dimension of a single stencil operatio…

cs.DB2025

Higher-Order Graph Databases

Maciej Besta, Shriram Chandran, Jakub Cudak +6

Recent advances in graph databases (GDBs) have been driving interest in large-scale analytics, yet current systems fail to support higher-order (HO) interactions beyond first-order…

cs.NI2024

PolarStar: Expanding the Scalability Horizon of Diameter-3 Networks

Kartik Lakhotia, Laura Monroe, Kelly Isham +4

We present PolarStar, a novel family of diameter-3 network topologies derived from the star product of low-diameter factor graphs. PolarStar gives the largest known diameter-3 netw…

cs.DC2017

sPIN: High-performance streaming Processing in the Network

Torsten Hoefler, Salvatore Di Girolamo, Konstantin Taranov +2

Optimizing communication performance is imperative for large-scale computing because communication overheads limit the strong scalability of parallel applications. Today's network…

cs.AI2025

Reasoning Language Models: A Blueprint

Maciej Besta, Julia Barth, Eric Schreiber +16

Reasoning language models (RLMs), also known as Large Reasoning Models (LRMs), such as OpenAI's o1 and o3, DeepSeek-R1, and Alibaba's QwQ, have redefined AI's problem-solving capab…

cs.AR2023

A High-performance, Energy-efficient Modular DMA Engine Architecture

Thomas Benz, Michael Rogenmoser, Paul Scheffler +5

Data transfers are essential in today's computing systems as latency and complex memory access patterns are increasingly challenging to manage. Direct memory access engines (DMAEs)…

cs.DC2026

PICO: Performance Insights for Collective Operations

Saverio Pasqualoni, Tommaso Bonato, Lorenzo Piarulli +3

Collective operations are cornerstones of both HPC applications and large-scale AI training and inference, yet benchmarking them in a systematic and reproducible way remains diffic…

cs.NI2026

The Multipath Reliable Connection (MRC) Transport

Rip Sohan, Eric Spada, Eric Davis +36

MRC is an open, production-grade transport designed for large-scale AI/ML training over best-effort Ethernet. It extends RoCEv2 with explicit, composable primitives for per-packet…

cs.DS2023

The spatial computer: A model for energy-efficient parallel computation

Lukas Gianinazzi, Tal Ben-Nun, Maciej Besta +4

We present a new parallel model of computation suitable for spatial architectures, for which the energy used for communication heavily depends on the distance of the communicating…

cs.DC2025

AI Factories: It's time to rethink the Cloud-HPC divide

Pedro Garcia Lopez, Daniel Barcelona Pons, Marcin Copik +7

The strategic importance of artificial intelligence is driving a global push toward Sovereign AI initiatives. Nationwide governments are increasingly developing dedicated infrastru…

cs.DC2023

FMI: Fast and Cheap Message Passing for Serverless Functions

Marcin Copik, Roman Böhringer, Alexandru Calotoiu +1

Serverless functions provide elastic scaling and a fine-grained billing model, making Function-as-a-Service (FaaS) an attractive programming model. However, for distributed jobs th…

cs.LG2024

EfQAT: An Efficient Framework for Quantization-Aware Training

Saleh Ashkboos, Bram Verhoef, Torsten Hoefler +2

Quantization-aware training (QAT) schemes have been shown to achieve near-full precision accuracy. They accomplish this by training a quantized model for multiple epochs. This is c…

cs.DC2026

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

Siyuan Shen, Anton Korzh, John Bachan +10

GPU collective communication is typically optimized for bandwidth, yet many emerging workloads are increasingly limited by latency. Long-context decode-heavy large language model (…

cs.CL2026

Large Language Model Selection with Limited Annotations

Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch +2

Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. T…

cs.DC2025

Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality

Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli +5

Communication locality plays a key role in the performance of collective operations on large HPC systems, especially on oversubscribed networks where groups of nodes are fully conn…

cs.DC2026

Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms

Zhiyi Hu, Siyuan Shen, Tommaso Bonato +6

The NVIDIA Collective Communication Library (NCCL) is a critical software layer enabling high-performance collectives on large-scale GPU clusters. Despite being open source with a…

cs.LG2019

Augment your batch: better training with larger batches

Elad Hoffer, Tal Ben-Nun, Itay Hubara +3

Large-batch SGD is important for scaling training of deep neural networks. However, without fine-tuning hyperparameter schedules, the generalization of the model may be hampered. W…

cs.DC2025

XaaS Containers: Performance-Portable Representation With Source and IR Containers

Marcin Copik, Eiman Alnuaimi, Alok Kamatar +6

High-performance computing (HPC) systems and cloud data centers are converging, and containers are becoming the default method of portable software deployment. Yet, while container…

cs.LG2022

Neural Parameter Allocation Search

Bryan A. Plummer, Nikoli Dryden, Julius Frost +2

Training neural networks requires increasing amounts of memory. Parameter sharing can reduce memory and communication costs, but existing methods assume networks have many identica…

cs.DC2017

Communication-Avoiding Parallel Algorithms for Solving Triangular Systems of Linear Equations

Tobias Wicky, Edgar Solomonik, Torsten Hoefler

We present a new parallel algorithm for solving triangular systems with multiple right hand sides (TRSM). TRSM is used extensively in numerical linear algebra computations, both to…

cs.LG2023

ASDL: A Unified Interface for Gradient Preconditioning in PyTorch

Kazuki Osawa, Satoki Ishikawa, Rio Yokota +2

Gradient preconditioning is a key technique to integrate the second-order information into gradients for improving and extending gradient-based learning algorithms. In deep learnin…

cs.DS2020

Parametric Graph Templates: Properties and Algorithms

Tal Ben-Nun, Lukas Gianinazzi, Torsten Hoefler +1

Hierarchical structure and repetition are prevalent in graphs originating from nature or engineering. These patterns can be represented by a class of parametric-structure graphs, w…

cs.AR2020

Slim NoC: A Low-Diameter On-Chip Network Topology for High Energy Efficiency and Scalability

Maciej Besta, Syed Minhaj Hassan, Sudhakar Yalamanchili +3

Emerging chips with hundreds and thousands of cores require networks with unprecedented energy/area efficiency and scalability. To address this, we propose Slim NoC (SN): a new on-…

cs.DC2024

FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an Example

Marcin Copik, Alexandru Calotoiu, Pengyu Zhou +2

FaaS (Function-as-a-Service) revolutionized cloud computing by replacing persistent virtual machines with dynamically allocated resources. This shift trades locality and statefulne…

cs.DC2024

Low-Depth Spatial Tree Algorithms

Yves Baumann, Tal Ben-Nun, Maciej Besta +3

Contemporary accelerator designs exhibit a high degree of spatial localization, wherein two-dimensional physical distance determines communication costs between processing elements…

cs.AR2020

Indirection Stream Semantic Register Architecture for Efficient Sparse-Dense Linear Algebra

Paul Scheffler, Florian Zaruba, Fabian Schuiki +2

Sparse-dense linear algebra is crucial in many domains, but challenging to handle efficiently on CPUs, GPUs, and accelerators alike; multiplications with sparse formats like CSR an…

cs.LG2025

BLaST: High Performance Inference and Pretraining using BLock Sparse Transformers

Patrik Okanovic, Sameer Deshmukh, Grzegorz Kwasniewski +8

The energy consumption of large-scale ML models is dominated by data movement, shuffling billions of parameters across memory hierarchies and data centers. Sparsification offers a…

cs.DC2024

High Performance Unstructured SpMM Computation Using Tensor Cores

Patrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini +3

High-performance sparse matrix-matrix (SpMM) multiplication is paramount for science and industry, as the ever-increasing sizes of data prohibit using dense data structures. Yet, e…

cs.DC2023

User-guided Page Merging for Memory Deduplication in Serverless Systems

Wei Qiu, Marcin Copik, Yun Wang +2

Serverless computing is an emerging cloud paradigm that offers an elastic and scalable allocation of computing resources with pay-as-you-go billing. In the Function-as-a-Service (F…

cs.LG2025

Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models

Wenqi Jiang, Marco Zeller, Roger Waleffe +2

A Retrieval-Augmented Language Model (RALM) combines a large language model (LLM) with a vector database to retrieve context-specific knowledge during text generation. This strateg…

cs.LG2018

μ-cuDNN: Accelerating Deep Learning Frameworks with Micro-Batching

Yosuke Oyama, Tal Ben-Nun, Torsten Hoefler +1

NVIDIA cuDNN is a low-level library that provides GPU kernels frequently used in deep learning. Specifically, cuDNN implements several equivalent convolution algorithms, whose perf…

cs.AR2024

LRSCwait: Enabling Scalable and Efficient Synchronization in Manycore Systems through Polling-Free and Retry-Free Operation

Samuel Riedel, Marc Gantenbein, Alessandro Ottaviano +2

Extensive polling in shared-memory manycore systems can lead to contention, decreased throughput, and poor energy efficiency. Both lock implementations and the general-purpose atom…

quant-ph2023

Disentangling Hype from Practicality: On Realistically Achieving Quantum Advantage

Torsten Hoefler, Thomas Haener, Matthias Troyer

Quantum computers offer a new paradigm of computing with the potential to vastly outperform any imagineable classical computer. This has caused a gold rush towards new quantum algo…

cs.DC2020

Enabling Highly-Scalable Remote Memory Access Programming with MPI-3 One Sided

Robert Gerstenberger, Maciej Besta, Torsten Hoefler

Modern interconnects offer remote direct memory access (RDMA) features. Yet, most applications rely on explicit message passing for communications albeit their unwanted overheads.…