activity
20172022
most citedThe gem5 Simulator: Version 20.0+

16 citations · 69 across the 23 of their papers we have counts for

collaborators
Showing cs.ARShow all

11 papers · 1 filter

cs.AR20223 cited

Enabling Flexibility for Sparse Tensor Acceleration via Heterogeneity

Eric Qin, Raveesh Garg, Abhimanyu Bambhaniya +5

Recently, numerous sparse hardware accelerators for Deep Neural Networks (DNNs), Graph Neural Networks (GNNs), and scientific computing applications have been proposed. A common ch…

cs.AR2021

RASA: Efficient Register-Aware Systolic Array Matrix Engine for CPU

Geonhwa Jeong, Eric Qin, Ananda Samajdar +4

As AI-based applications become pervasive, CPU vendors are starting to incorporate matrix engines within the datapath to boost efficiency. Systolic arrays have been the premier arc…

cs.AR2021

Union: A Unified HW-SW Co-Design Ecosystem in MLIR for Evaluating Tensor Operations on Spatial Accelerators

Geonhwa Jeong, Gokcen Kestor, Prasanth Chatarasi +5

To meet the extreme compute demands for deep learning across commercial and scientific applications, dataflow accelerators are becoming increasingly popular. While these "domain-sp…

cs.AR2020

Architecture, Dataflow and Physical Design Implications of 3D-ICs for DNN-Accelerators

Jan Moritz Joseph, Ananda Samajdar, Lingjun Zhu +4

The everlasting demand for higher computing power for deep neural networks (DNNs) drives the development of parallel computing architectures. 3D integration, in which chips are int…

cs.AR2020

Dataflow-Architecture Co-Design for 2.5D DNN Accelerators using Wireless Network-on-Package

Robert Guirado, Hyoukjun Kwon, Sergi Abadal +2

Deep neural network (DNN) models continue to grow in size and complexity, demanding higher computational power to enable real-time inference. To efficiently deliver such computatio…

cs.AR202010 cited

ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning

Sheng-Chun Kao, Geonhwa Jeong, Tushar Krishna

DNN accelerators provide efficiency by leveraging reuse of activations/weights/outputs during the DNN computations to reduce data movement from DRAM to the chip. The reuse is captu…