activity
20182022
most citedServing Recurrent Neural Networks Efficiently with a Spatial Accelerator

10 citations · 15 across the 3 of their papers we have counts for

collaborators

5 papers

cs.PL20224 cited

Stardust: Compiling Sparse Tensor Algebra to a Reconfigurable Dataflow Architecture

Olivia Hsu, Alexander Rucker, Tian Zhao +2

We introduce Stardust, a compiler that compiles sparse tensor algebra to reconfigurable dataflow architectures (RDAs). Stardust introduces new user-provided data representation and…

cs.AR20221 cited

Efficient Memory Partitioning in Software Defined Hardware

Matthew Feldman, Tian Zhao, Kunle Olukotun

As programmers turn to software-defined hardware (SDH) to maintain a high level of productivity while programming hardware to run complex algorithms, heavy-lifting must be done by…

cs.AR2021

Capstan: A Vector RDA for Sparsity

Alexander Rucker, Matthew Vilim, Tian Zhao +3

This paper proposes Capstan: a scalable, parallel-patterns-based, reconfigurable dataflow accelerator (RDA) for sparse and dense tensor applications. Instead of designing for one a…

cs.DC201910 cited

Serving Recurrent Neural Networks Efficiently with a Spatial Accelerator

Tian Zhao, Yaqi Zhang, Kunle Olukotun

Recurrent Neural Network (RNN) applications form a major class of AI-powered, low-latency data center workloads. Most execution models for RNN acceleration break computation graphs…

cs.LG2018

Analysis of DAWNBench, a Time-to-Accuracy Machine Learning Performance Benchmark

Cody Coleman, Daniel Kang, Deepak Narayanan +7

Researchers have proposed hardware, software, and algorithmic optimizations to improve the computational performance of deep learning. While some of these optimizations perform the…