Publications (260)
Ultra Ethernet's Design Principles and Architectural Innovations
Torsten Hoefler, Karen Schramm, Eric Spada +13
The recently released Ultra Ethernet (UE) 1.0 specification defines a transformative High-Performance Ethernet standard for future Artificial Intelligence (AI) and High-Performance…
Benchmarking Filtered Approximate Nearest Neighbor Search Algorithms on Transformer-based Embedding Vectors
Patrick Iff, Paul Bruegger, Marcin Chrapek +3
Advances in embedding models for text, image, audio, and video drive progress across multiple domains, including retrieval-augmented generation, recommendation systems, and others.…
Swing: Short-cutting Rings for Higher Bandwidth Allreduce
Daniele De Sensi, Tommaso Bonato, David Saam +1
The allreduce collective operation accounts for a significant fraction of the runtime of workloads running on distributed systems. One factor determining its performance is the dis…
GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding
Jiale Chen, Torsten Hoefler, Dan Alistarh
The paper introduces GPTQ-2D, an algorithm that rounds a real matrix to integers under a two-sided quadratic metric in cubic time by processing entries anti-diagonal by anti-diagon…
Cached Operator Reordering: A Unified View for Fast GNN Training
Julia Bazinska, Andrei Ivanov, Tal Ben-Nun +4
Graph Neural Networks (GNNs) are a powerful tool for handling structured graph data and addressing tasks such as node classification, graph classification, and clustering. However,…
Stateful Dataflow Multigraphs: A Data-Centric Model for Performance Portability on Heterogeneous Architectures
Tal Ben-Nun, Johannes de Fine Licht, Alexandros Nikolaos Ziogas +2
The ubiquity of accelerators in high-performance computing has driven programming complexity beyond the skill-set of the average domain scientist. To maintain performance portabili…
Demystifying NVSHMEM: A System-Level Analysis on Symmetric Memory and Device-Initiated Operations in GPU Communication
Yijun Ma, Siyuan Shen, Tiancheng Chen +6
NVSHMEM is NVIDIA's OpenSHMEM-based PGAS communication library for GPU clusters, enabling GPU-initiated, one-sided communication through symmetric memory. Despite its growing adopt…
Ab-initio Quantum Transport with the GW Approximation, 42,240 Atoms, and Sustained Exascale Performance
Nicolas Vetsch, Alexander Maeder, Vincent Maillou +7
Designing nanoscale electronic devices such as the currently manufactured nanoribbon field-effect transistors (NRFETs) requires advanced modeling tools capturing all relevant quant…
RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures
Patrick Iff, Benigna Bruggmann, Blaise Morel +3
Chiplet architectures are on the rise as they promise to overcome the scaling challenges of monolithic chips. A key component of such architectures is an efficient inter-chiplet in…
Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun +2
The growing energy and performance costs of deep learning have driven the community to reduce the size of neural networks by selectively pruning components. Similarly to their biol…
A scalable weakly-synchronous algorithm for solving partial differential equations
Konduri Aditya, Tobias Gysi, Grzegorz Kwasniewski +3
Synchronization overheads pose a major challenge as applications advance towards extreme scales. In current large-scale algorithms, synchronization as well as data communication de…
Flare: Flexible In-Network Allreduce
Daniele De Sensi, Salvatore Di Girolamo, Saleh Ashkboos +2
The allreduce operation is one of the most commonly used communication routines in distributed applications. To improve its bandwidth and to reduce network traffic, this operation…
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
Mikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek +3
In the Fully Sharded Data Parallel (FSDP) training pipeline, collective operations can be interleaved to maximize the communication/computation overlap. In this scenario, outstandi…
Extracting Clean Performance Models from Tainted Programs
Marcin Copik, Alexandru Calotoiu, Tobias Grosser +3
Performance models are well-known instruments to understand the scaling behavior of parallel applications. They express how performance changes as key execution parameters, such as…
Demystifying Graph Databases: Analysis and Taxonomy of Data Organization, System Designs, and Graph Queries
Maciej Besta, Robert Gerstenberger, Emanuel Peter +5
Graph processing has become an important part of multiple areas of computer science, such as machine learning, computational sciences, medical applications, social network analysis…
Slim Graph: Practical Lossy Graph Compression for Approximate Graph Processing, Storage, and Analytics
Maciej Besta, Simon Weber, Lukas Gianinazzi +4
We propose Slim Graph: the first programming model and framework for practical lossy graph compression that facilitates high-performance approximate graph processing, storage, and…
A Data-Centric Approach to Extreme-Scale Ab initio Dissipative Quantum Transport Simulations
Alexandros Nikolaos Ziogas, Tal Ben-Nun, Guillermo Indalecio Fernández +3
The computational efficiency of a state of the art ab initio quantum transport (QT) solver, capable of revealing the coupled electro-thermal properties of atomically-resolved nano-…
GraphSeek: Next-Generation Graph Analytics with LLMs
Maciej Besta, Åukasz Jarmocik, Orest Hrycyna +7
Graphs are foundational across domains but remain hard to use without deep expertise. LLMs promise accessible natural language (NL) graph analytics, yet they fail to process indust…
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci +6
We introduce QuaRot, a new Quantization scheme based on Rotations, which is able to quantize LLMs end-to-end, including all weights, activations, and KV cache in 4 bits. QuaRot rot…
Demystifying Higher-Order Graph Neural Networks
Maciej Besta, Florian Scheidl, Lukas Gianinazzi +4
Higher-order graph neural networks (HOGNNs) and the related architectures from Topological Deep Learning are an important class of GNN models that harness polyadic relations betwee…
Efficient Quantized Sparse Matrix Operations on Tensor Cores
Shigang Li, Kazuki Osawa, Torsten Hoefler
The exponentially growing model size drives the continued success of deep learning, but it brings prohibitive computation and memory cost. From the algorithm perspective, model spa…
PsPIN: A high-performance low-power architecture for flexible in-network compute
Salvatore Di Girolamo, Andreas Kurth, Alexandru Calotoiu +5
The capacity of offloading data and control tasks to the network is becoming increasingly important, especially if we consider the faster growth of network speed when compared to C…
Fortify Your Foundations: Practical Privacy and Security for Foundation Model Deployments In The Cloud
Marcin Chrapek, Anjo Vahldiek-Oberwagner, Marcin Spoczynski +3
Foundation Models (FMs) display exceptional performance in tasks such as natural language processing and are being applied across a growing range of disciplines. Although typically…
Parallel Algorithms for Finding Large Cliques in Sparse Graphs
Lukas Gianinazzi, Maciej Besta, Yannick Schaffner +1
We present a parallel k-clique listing algorithm with improved work bounds (for the same depth) in sparse graphs with low degeneracy or arboricity. We achieve this by introducing a…
Sparse Stream Semantic Registers: A Lightweight ISA Extension Accelerating General Sparse Linear Algebra
Paul Scheffler, Florian Zaruba, Fabian Schuiki +2
Sparse linear algebra is crucial in many application domains, but challenging to handle efficiently in both software and hardware, with one- and two-sided operand sparsity handled…
XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing
Torsten Hoefler, Marcin Copik, Pete Beckman +8
HPC and Cloud have evolved independently, specializing their innovations into performance or productivity. Acceleration as a Service (XaaS) is a recipe to empower both fields with…
All models are wrong, some are useful: Model Selection with Limited Labels
Patrik Okanovic, Andreas Kirsch, Jannes Kasper +3
We introduce MODEL SELECTOR, a framework for label-efficient selection of pretrained classifiers. Given a pool of unlabeled target data, MODEL SELECTOR samples a small subset of hi…
Practice of Streaming Processing of Dynamic Graphs: Concepts, Models, and Systems
Maciej Besta, Marc Fischer, Vasiliki Kalavri +2
Graph processing has become an important part of various areas of computing, including machine learning, medical applications, social network analysis, computational sciences, and…
Deinsum: Practically I/O Optimal Multilinear Algebra
Alexandros Nikolaos Ziogas, Grzegorz Kwasniewski, Tal Ben-Nun +2
Multilinear algebra kernel performance on modern massively-parallel systems is determined mainly by data movement. However, deriving data movement-optimal distributed schedules for…
Streaming Message Interface: High-Performance Distributed Memory Programming on Reconfigurable Hardware
Tiziano De Matteis, Johannes de Fine Licht, Jakub Beránek +1
Distributed memory programming is the established paradigm used in high-performance computing (HPC) systems, requiring explicit communication between nodes and devices. When FPGAs…
HexaMesh: Scaling to Hundreds of Chiplets with an Optimized Chiplet Arrangement
Patrick Iff, Maciej Besta, Matheus Cavalcante +3
2.5D integration is an important technique to tackle the growing cost of manufacturing chips in advanced technology nodes. This poses the challenge of providing high-performance in…
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
Siyuan Shen, Mikhail Khalilov, Lukas Gianinazzi +6
Resource disaggregation is a promising technique for improving the efficiency of large-scale computing systems. However, this comes at the cost of increased memory access latency d…
DiffDA: a Diffusion Model for Weather-scale Data Assimilation
Langwen Huang, Lukas Gianinazzi, Yuejiang Yu +2
The generation of initial conditions via accurate data assimilation is crucial for weather forecasting and climate modeling. We propose DiffDA as a denoising diffusion model capabl…
The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm
Jiale Chen, Yalda Shabanzadeh, Elvir CrnÄeviÄ +2
Quantizing the weights of large language models (LLMs) from 16-bit to lower bitwidth is the de facto approach to deploy massive transformers onto more affordable accelerators. Whil…
DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing
Afif Boudaoud, Alexandru Calotoiu, Marcin Copik +1
Automatic differentiation (AD) is a set of techniques that systematically applies the chain rule to compute the gradients of functions without requiring human intervention. Althoug…
Core Hours and Carbon Credits: Incentivizing Sustainability in HPC
Alok Kamatar, Maxime Gonthier, Valerie Hayot-Sasson +6
Realizing a shared responsibility between providers and consumers is critical to manage the sustainability of HPC. However, while cost may motivate efficiency improvements by infra…
Error bounded compression for weather and climate applications
Langwen Huang, Luigi Fusco, Florian Scheidl +4
As the resolution of weather and climate simulations increases, the amount of data produced is growing rapidly from hundreds of terabytes to tens of petabytes. The huge size become…
Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal
Wenqi Jiang, Hang Hu, Torsten Hoefler +1
Vector search systems are indispensable in large language model (LLM) serving, search engines, and recommender systems, where minimizing online search latency is essential. Among v…
MLIR-Forge: A Modular Framework for Language Smiths
Berke Ates, Philipp Schaad, Timo Schneider +2
Optimizing compilers are essential for the efficient and correct execution of software across various scientific fields. Domain-specific languages (DSL) typically use higher level…
The Convergence of Sparsified Gradient Methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson +3
Distributed training of massive machine learning models, in particular deep neural networks, via Stochastic Gradient Descent (SGD) is becoming commonplace. Several families of comm…
A High-Performance Design, Implementation, Deployment, and Evaluation of The Slim Fly Network
Nils Blach, Maciej Besta, Daniele De Sensi +10
Novel low-diameter network topologies such as Slim Fly (SF) offer significant cost and power advantages over the established Fat Tree, Clos, or Dragonfly. To spearhead the adoption…
Active Access: A Mechanism for High-Performance Distributed Data-Centric Computations
Maciej Besta, Torsten Hoefler
Remote memory access (RMA) is an emerging high-performance programming model that uses RDMA hardware directly. Yet, accessing remote memories cannot invoke activities at the target…
SpaDA: A Spatial Dataflow Architecture Programming Language
Lukas Gianinazzi, Tal Ben-Nun, Torsten Hoefler
Spatial dataflow architectures like the Cerebras Wafer-Scale Engine deliver exceptional performance in AI and scientific computing by distributing scratchpad memory across hundreds…
Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge
Drago Plecko, Patrik Okanovic, Shreyas Havaldar +2
Artificial intelligence (AI) systems hold great promise for advancing various scientific disciplines, and are increasingly used in real-world applications. Despite their remarkable…
MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
Elias Frantar, Roberto L. Castro, Jiale Chen +2
As inference on Large Language Models (LLMs) emerges as an important workload in machine learning applications, weight quantization has become a standard technique for efficient GP…
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
Siyuan Shen, Langwen Huang, Marcin Chrapek +5
The shift towards high-bandwidth networks driven by AI workloads in data centers and HPC clusters has unintentionally aggravated network latency, adversely affecting the performanc…
Evolving HPC services to enable ML workloads on HPE Cray EX
Stefano Schuppli, Fawzi Mohamed, Henrique Mendonça +10
The Alps Research Infrastructure leverages GH200 technology at scale, featuring 10,752 GPUs. Accessing Alps provides a significant computational advantage for researchers in Artifi…
A Data-Centric Optimization Framework for Machine Learning
Oliver Rausch, Tal Ben-Nun, Nikoli Dryden +3
Rapid progress in deep learning is leading to a diverse set of quickly changing models, with a dramatically growing demand for compute. However, as frameworks specialize performanc…
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
Timo Schneider, Pengcheng Xu, Torsten Hoefler
In the era of post-Moore computing, network offload emerges as a solution to two challenges: the imperative for low-latency communication and the push towards hardware specialisati…
Communication Lower Bounds of Bilinear Algorithms for Symmetric Tensor Contractions
Edgar Solomonik, James Demmel, Torsten Hoefler
We introduce a new theoretical framework for deriving lower bounds on data movement in bilinear algorithms. Bilinear algorithms are a general representation of fast algorithms for…
Optimizing the Data Movement in Quantum Transport Simulations via Data-Centric Parallel Programming
Alexandros Nikolaos Ziogas, Tal Ben-Nun, Guillermo Indalecio Fernández +3
Designing efficient cooling systems for integrated circuits (ICs) relies on a deep understanding of the electro-thermal properties of transistors. To shed light on this issue in cu…
ScalAna: Automating Scaling Loss Detection with Graph Analysis
Yuyang Jin, Haojie Wang, Teng Yu +4
Scaling a parallel program to modern supercomputers is challenging due to inter-process communication, Amdahl's law, and resource contention. Performance analysis tools for finding…
Denoising Application Performance Models with Noise-Resilient Priors
Gustavo de Morais, Alexander GeiÃ, Alexandru Calotoiu +5
As parallel codes are scaled to larger computing systems, performance models play a crucial role in identifying potential bottlenecks. However, constructing these models analytical…
Near-Optimal Wafer-Scale Reduce
Piotr Luczynski, Lukas Gianinazzi, Patrick Iff +3
Efficient Reduce and AllReduce communication collectives are a critical cornerstone of high-performance computing (HPC) applications. We present the first systematic investigation…
Temporal Vectorization: A Compiler Approach to Automatic Multi-Pumping
Carl-Johannes Johnsen, Tiziano De Matteis, Tal Ben-Nun +2
The multi-pumping resource sharing technique can overcome the limitations commonly found in single-clocked FPGA designs by allowing hardware components to operate at a higher clock…
Spatial Mixture-of-Experts
Nikoli Dryden, Torsten Hoefler
Many data have an underlying dependence on spatial location; it may be weather on the Earth, a simulation on a mesh, or a registered image. Yet this feature is rarely taken advanta…
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
Aofeng Shen, Chi Zhang, Yakup Budanaz +4
Tile-based many-Processing Element (PE) accelerators can achieve competitive performance on General Matrix Multiplication (GEMM), but they are extremely hard to program, as their o…
When Data Is Scarce: Scaling Sparse Language Models with Repeated Training
Boqian Wu, Qiao Xiao, Patrik Okanovic +6
Scaling laws for dense LLMs under infinite data are well explored, but how sparsity interacts with limited data is not. In this work, we study sparse training in data-constrained r…
Motif Prediction with Graph Neural Networks
Maciej Besta, Raphael Grob, Cesare Miglioli +8
Link prediction is one of the central problems in graph mining. However, recent studies highlight the importance of higher-order network analysis, where complex structures called m…
Hardware Acceleration for Knowledge Graph Processing: Challenges & Recent Developments
Maciej Besta, Robert Gerstenberger, Patrick Iff +9
Knowledge graphs (KGs) have achieved significant attention in recent years, particularly in the area of the Semantic Web as well as gaining popularity in other application domains…
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
Luigi Fusco, Mikhail Khalilov, Marcin Chrapek +3
Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads…
REPS: Recycled Entropy Packet Spraying for Adaptive Load Balancing and Failure Mitigation
Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini +7
Next-generation datacenters require highly efficient network load balancing to manage the growing scale of artificial intelligence (AI) training and general datacenter traffic. How…
OSMOSIS: Enabling Multi-Tenancy in Datacenter SmartNICs
Mikhail Khalilov, Marcin Chrapek, Siyuan Shen +7
Multi-tenancy is essential for unleashing SmartNIC's potential in datacenters. Our systematic analysis in this work shows that existing on-path SmartNICs have resource multiplexing…
Flexible Communication Avoiding Matrix Multiplication on FPGA with High-Level Synthesis
Johannes de Fine Licht, Grzegorz Kwasniewski, Torsten Hoefler
Data movement is the dominating factor affecting performance and energy in modern computing systems. Consequently, many algorithms have been developed to minimize the number of I/O…
A Fast Analytical Model of Fully Associative Caches
Tobias Gysi, Tobias Grosser, Laurin Brandner +1
While the cost of computation is an easy to understand local property, the cost of data movement on cached architectures depends on global state, does not compose, and is hard to p…
Stream Semantic Registers: A Lightweight RISC-V ISA Extension Achieving Full Compute Utilization in Single-Issue Cores
Fabian Schuiki, Florian Zaruba, Torsten Hoefler +1
Single-issue processor cores are very energy efficient but suffer from the von Neumann bottleneck, in that they must explicitly fetch and issue the loads/storse necessary to feed t…
StencilFlow: Mapping Large Stencil Programs to Distributed Spatial Computing Systems
Johannes de Fine Licht, Andreas Kuster, Tiziano De Matteis +3
Spatial computing devices have been shown to significantly accelerate stencil computations, but have so far relied on unrolling the iterative dimension of a single stencil operatio…
Higher-Order Graph Databases
Maciej Besta, Shriram Chandran, Jakub Cudak +6
Recent advances in graph databases (GDBs) have been driving interest in large-scale analytics, yet current systems fail to support higher-order (HO) interactions beyond first-order…
PolarStar: Expanding the Scalability Horizon of Diameter-3 Networks
Kartik Lakhotia, Laura Monroe, Kelly Isham +4
We present PolarStar, a novel family of diameter-3 network topologies derived from the star product of low-diameter factor graphs. PolarStar gives the largest known diameter-3 netw…
sPIN: High-performance streaming Processing in the Network
Torsten Hoefler, Salvatore Di Girolamo, Konstantin Taranov +2
Optimizing communication performance is imperative for large-scale computing because communication overheads limit the strong scalability of parallel applications. Today's network…
Reasoning Language Models: A Blueprint
Maciej Besta, Julia Barth, Eric Schreiber +16
Reasoning language models (RLMs), also known as Large Reasoning Models (LRMs), such as OpenAI's o1 and o3, DeepSeek-R1, and Alibaba's QwQ, have redefined AI's problem-solving capab…
A High-performance, Energy-efficient Modular DMA Engine Architecture
Thomas Benz, Michael Rogenmoser, Paul Scheffler +5
Data transfers are essential in today's computing systems as latency and complex memory access patterns are increasingly challenging to manage. Direct memory access engines (DMAEs)…
PICO: Performance Insights for Collective Operations
Saverio Pasqualoni, Tommaso Bonato, Lorenzo Piarulli +3
Collective operations are cornerstones of both HPC applications and large-scale AI training and inference, yet benchmarking them in a systematic and reproducible way remains diffic…
The Multipath Reliable Connection (MRC) Transport
Rip Sohan, Eric Spada, Eric Davis +36
MRC is an open, production-grade transport designed for large-scale AI/ML training over best-effort Ethernet. It extends RoCEv2 with explicit, composable primitives for per-packet…
The spatial computer: A model for energy-efficient parallel computation
Lukas Gianinazzi, Tal Ben-Nun, Maciej Besta +4
We present a new parallel model of computation suitable for spatial architectures, for which the energy used for communication heavily depends on the distance of the communicating…
AI Factories: It's time to rethink the Cloud-HPC divide
Pedro Garcia Lopez, Daniel Barcelona Pons, Marcin Copik +7
The strategic importance of artificial intelligence is driving a global push toward Sovereign AI initiatives. Nationwide governments are increasingly developing dedicated infrastru…
FMI: Fast and Cheap Message Passing for Serverless Functions
Marcin Copik, Roman Böhringer, Alexandru Calotoiu +1
Serverless functions provide elastic scaling and a fine-grained billing model, making Function-as-a-Service (FaaS) an attractive programming model. However, for distributed jobs th…
EfQAT: An Efficient Framework for Quantization-Aware Training
Saleh Ashkboos, Bram Verhoef, Torsten Hoefler +2
Quantization-aware training (QAT) schemes have been shown to achieve near-full precision accuracy. They accomplish this by training a quantized model for multiple epochs. This is c…
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives
Siyuan Shen, Anton Korzh, John Bachan +10
GPU collective communication is typically optimized for bandwidth, yet many emerging workloads are increasingly limited by latency. Long-context decode-heavy large language model (…
Large Language Model Selection with Limited Annotations
Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch +2
Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. T…
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli +5
Communication locality plays a key role in the performance of collective operations on large HPC systems, especially on oversubscribed networks where groups of nodes are fully conn…
Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms
Zhiyi Hu, Siyuan Shen, Tommaso Bonato +6
The NVIDIA Collective Communication Library (NCCL) is a critical software layer enabling high-performance collectives on large-scale GPU clusters. Despite being open source with a…
Augment your batch: better training with larger batches
Elad Hoffer, Tal Ben-Nun, Itay Hubara +3
Large-batch SGD is important for scaling training of deep neural networks. However, without fine-tuning hyperparameter schedules, the generalization of the model may be hampered. W…
XaaS Containers: Performance-Portable Representation With Source and IR Containers
Marcin Copik, Eiman Alnuaimi, Alok Kamatar +6
High-performance computing (HPC) systems and cloud data centers are converging, and containers are becoming the default method of portable software deployment. Yet, while container…
Neural Parameter Allocation Search
Bryan A. Plummer, Nikoli Dryden, Julius Frost +2
Training neural networks requires increasing amounts of memory. Parameter sharing can reduce memory and communication costs, but existing methods assume networks have many identica…
Communication-Avoiding Parallel Algorithms for Solving Triangular Systems of Linear Equations
Tobias Wicky, Edgar Solomonik, Torsten Hoefler
We present a new parallel algorithm for solving triangular systems with multiple right hand sides (TRSM). TRSM is used extensively in numerical linear algebra computations, both to…
ASDL: A Unified Interface for Gradient Preconditioning in PyTorch
Kazuki Osawa, Satoki Ishikawa, Rio Yokota +2
Gradient preconditioning is a key technique to integrate the second-order information into gradients for improving and extending gradient-based learning algorithms. In deep learnin…
Parametric Graph Templates: Properties and Algorithms
Tal Ben-Nun, Lukas Gianinazzi, Torsten Hoefler +1
Hierarchical structure and repetition are prevalent in graphs originating from nature or engineering. These patterns can be represented by a class of parametric-structure graphs, w…
Slim NoC: A Low-Diameter On-Chip Network Topology for High Energy Efficiency and Scalability
Maciej Besta, Syed Minhaj Hassan, Sudhakar Yalamanchili +3
Emerging chips with hundreds and thousands of cores require networks with unprecedented energy/area efficiency and scalability. To address this, we propose Slim NoC (SN): a new on-…
FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an Example
Marcin Copik, Alexandru Calotoiu, Pengyu Zhou +2
FaaS (Function-as-a-Service) revolutionized cloud computing by replacing persistent virtual machines with dynamically allocated resources. This shift trades locality and statefulne…
Low-Depth Spatial Tree Algorithms
Yves Baumann, Tal Ben-Nun, Maciej Besta +3
Contemporary accelerator designs exhibit a high degree of spatial localization, wherein two-dimensional physical distance determines communication costs between processing elements…
Indirection Stream Semantic Register Architecture for Efficient Sparse-Dense Linear Algebra
Paul Scheffler, Florian Zaruba, Fabian Schuiki +2
Sparse-dense linear algebra is crucial in many domains, but challenging to handle efficiently on CPUs, GPUs, and accelerators alike; multiplications with sparse formats like CSR an…
BLaST: High Performance Inference and Pretraining using BLock Sparse Transformers
Patrik Okanovic, Sameer Deshmukh, Grzegorz Kwasniewski +8
The energy consumption of large-scale ML models is dominated by data movement, shuffling billions of parameters across memory hierarchies and data centers. Sparsification offers a…
High Performance Unstructured SpMM Computation Using Tensor Cores
Patrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini +3
High-performance sparse matrix-matrix (SpMM) multiplication is paramount for science and industry, as the ever-increasing sizes of data prohibit using dense data structures. Yet, e…
User-guided Page Merging for Memory Deduplication in Serverless Systems
Wei Qiu, Marcin Copik, Yun Wang +2
Serverless computing is an emerging cloud paradigm that offers an elastic and scalable allocation of computing resources with pay-as-you-go billing. In the Function-as-a-Service (F…
Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models
Wenqi Jiang, Marco Zeller, Roger Waleffe +2
A Retrieval-Augmented Language Model (RALM) combines a large language model (LLM) with a vector database to retrieve context-specific knowledge during text generation. This strateg…
μ-cuDNN: Accelerating Deep Learning Frameworks with Micro-Batching
Yosuke Oyama, Tal Ben-Nun, Torsten Hoefler +1
NVIDIA cuDNN is a low-level library that provides GPU kernels frequently used in deep learning. Specifically, cuDNN implements several equivalent convolution algorithms, whose perf…
LRSCwait: Enabling Scalable and Efficient Synchronization in Manycore Systems through Polling-Free and Retry-Free Operation
Samuel Riedel, Marc Gantenbein, Alessandro Ottaviano +2
Extensive polling in shared-memory manycore systems can lead to contention, decreased throughput, and poor energy efficiency. Both lock implementations and the general-purpose atom…
Disentangling Hype from Practicality: On Realistically Achieving Quantum Advantage
Torsten Hoefler, Thomas Haener, Matthias Troyer
Quantum computers offer a new paradigm of computing with the potential to vastly outperform any imagineable classical computer. This has caused a gold rush towards new quantum algo…
Enabling Highly-Scalable Remote Memory Access Programming with MPI-3 One Sided
Robert Gerstenberger, Maciej Besta, Torsten Hoefler
Modern interconnects offer remote direct memory access (RDMA) features. Yet, most applications rely on explicit message passing for communications albeit their unwanted overheads.…