activity
20172023
most citedASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale

78 citations · 171 across the 33 of their papers we have counts for

collaborators
Showing 2021 · cs.DCShow all

5 papers · 2 filters

cs.DC2021

Themis: A Network Bandwidth-Aware Collective Scheduling Policy for Distributed Training of DL Models

Saeed Rashidi, William Won, Sudarshan Srinivasan +2

Distributed training is a solution to reduce DNN training time by splitting the task across multiple NPUs (e.g., GPU/TPU). However, distributed training adds communication overhead…

cs.DC2021

LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models

William Won, Saeed Rashidi, Sudarshan Srinivasan +1

As model sizes in machine learning continue to scale, distributed training is necessary to accommodate model weights within each device and to reduce training time. However, this c…

cs.DC2021

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

Gordon E. Moon, Hyoukjun Kwon, Geonhwa Jeong +3

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via…

cs.DC2021

Extending Sparse Tensor Accelerators to Support Multiple Compression Formats

Eric Qin, Geonhwa Jeong, William Won +7

Sparsity, which occurs in both scientific applications and Deep Learning (DL) models, has been a key target of optimization within recent ASIC accelerators due to the potential mem…

cs.DC2021

Understanding the Design-Space of Sparse/Dense Multiphase GNN dataflows on Spatial Accelerators

Raveesh Garg, Eric Qin, Francisco Muñoz-Martínez +8

Graph Neural Networks (GNNs) have garnered a lot of recent interest because of their success in learning representations from graph-structured data across several critical applicat…