activity
20172023
most citedASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale

78 citations · 222 across the 36 of their papers we have counts for

collaborators
Showing 2020 · cs.ARShow all

6 papers · 2 filters

cs.AR2020

Architecture, Dataflow and Physical Design Implications of 3D-ICs for DNN-Accelerators

Jan Moritz Joseph, Ananda Samajdar, Lingjun Zhu +4

The everlasting demand for higher computing power for deep neural networks (DNNs) drives the development of parallel computing architectures. 3D integration, in which chips are int…

cs.AR2020

Dataflow-Architecture Co-Design for 2.5D DNN Accelerators using Wireless Network-on-Package

Robert Guirado, Hyoukjun Kwon, Sergi Abadal +2

Deep neural network (DNN) models continue to grow in size and complexity, demanding higher computational power to enable real-time inference. To efficiently deliver such computatio…

cs.AR2020★ 10 cited

ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning

Sheng-Chun Kao, Geonhwa Jeong, Tushar Krishna

DNN accelerators provide efficiency by leveraging reuse of activations/weights/outputs during the DNN computations to reduce data movement from DRAM to the chip. The reuse is captu…

cs.AR2020

Breaking Barriers: Maximizing Array Utilization for Compute In-Memory Fabrics

Brian Crafton, Samuel Spetalnick, Gauthaman Murali +3

Compute in-memory (CIM) is a promising technique that minimizes data transport, the primary performance bottleneck and energy cost of most data intensive applications. This has fou…

cs.AR2020★ 16 cited

The gem5 Simulator: Version 20.0+

Jason Lowe-Power, Abdul Mutaal Ahmad, Ayaz Akram +75

The open-source and community-supported gem5 simulator is one of the most popular tools for computer architecture research. This simulation infrastructure allows researchers to mod…

cs.AR2020

Enabling Compute-Communication Overlap in Distributed Deep Learning Training Platforms

Saeed Rashidi, Matthew Denton, Srinivas Sridharan +4

Deep Learning (DL) training platforms are built by interconnecting multiple DL accelerators (e.g., GPU/TPU) via fast, customized interconnects with 100s of gigabytes (GBs) of bandw…