1 citations · 2 across the 3 of their papers we have counts for
5 papers
S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN Acceleration
Zhi-Gang Liu, Paul N. Whatmough, Yuhao Zhu +1
Exploiting sparsity is a key technique in accelerating quantized convolutional neural network (CNN) inference on mobile devices. Prior sparse CNN accelerators largely exploit un-st…
Doping: A technique for efficient compression of LSTM models using sparse structured additive matrices
Urmish Thakker, Paul N. Whatmough, Zhigang Liu +2
Structured matrices, such as those derived from Kronecker products (KP), are effective at compressing neural networks, but can lead to unacceptable accuracy loss when applied to la…
Sparse Systolic Tensor Array for Efficient CNN Hardware Acceleration
Zhi-Gang Liu, Paul N. Whatmough, Matthew Mattina
Convolutional neural network (CNN) inference on mobile devices demands efficient hardware acceleration of low-precision (INT8) general matrix multiplication (GEMM). Exploiting data…
Efficient Residue Number System Based Winograd Convolution
Zhi-Gang Liu, Matthew Mattina
Prior research has shown that Winograd algorithm can reduce the computational complexity of convolutional neural networks (CNN) with weights and activations represented in floating…
Systolic Tensor Array: An Efficient Structured-Sparse GEMM Accelerator for Mobile CNN Inference
Zhi-Gang Liu, Paul N. Whatmough, Matthew Mattina
Convolutional neural network (CNN) inference on mobile devices demands efficient hardware acceleration of low-precision (INT8) general matrix multiplication (GEMM). The systolic ar…