most citedSystolic Tensor Array: An Efficient Structured-Sparse GEMM Accelerator for Mobile CNN Inference

1 citations · 2 across the 3 of their papers we have counts for

collaborators

5 papers

cs.AR2021

S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN Acceleration

Zhi-Gang Liu, Paul N. Whatmough, Yuhao Zhu +1

Exploiting sparsity is a key technique in accelerating quantized convolutional neural network (CNN) inference on mobile devices. Prior sparse CNN accelerators largely exploit un-st…

cs.LG2021★ 1 cited

Doping: A technique for efficient compression of LSTM models using sparse structured additive matrices

Urmish Thakker, Paul N. Whatmough, Zhigang Liu +2

Structured matrices, such as those derived from Kronecker products (KP), are effective at compressing neural networks, but can lead to unacceptable accuracy loss when applied to la…

cs.AR2020

Sparse Systolic Tensor Array for Efficient CNN Hardware Acceleration

Zhi-Gang Liu, Paul N. Whatmough, Matthew Mattina

Convolutional neural network (CNN) inference on mobile devices demands efficient hardware acceleration of low-precision (INT8) general matrix multiplication (GEMM). Exploiting data…

cs.LG2020

Efficient Residue Number System Based Winograd Convolution

Zhi-Gang Liu, Matthew Mattina

Prior research has shown that Winograd algorithm can reduce the computational complexity of convolutional neural networks (CNN) with weights and activations represented in floating…

cs.DC2020★ 1 cited

Systolic Tensor Array: An Efficient Structured-Sparse GEMM Accelerator for Mobile CNN Inference

Zhi-Gang Liu, Paul N. Whatmough, Matthew Mattina

Convolutional neural network (CNN) inference on mobile devices demands efficient hardware acceleration of low-precision (INT8) general matrix multiplication (GEMM). The systolic ar…