2 papers
cs.LG2020
GrateTile: Efficient Sparse Tensor Tiling for CNN Processing
Yu-Sheng Lin, Hung Chang Lu, Yang-Bin Tsao +3
We propose GrateTile, an efficient, hardwarefriendly data storage scheme for sparse CNN feature maps (activations). It divides data into uneven-sized subtensors and, with small ind…
cs.DC2019
MERIT: Tensor Transform for Memory-Efficient Vision Processing on Parallel Architectures
Yu-Sheng Lin, Wei-Chao Chen, Shao-Yi Chien
Computationally intensive deep neural networks (DNNs) are well-suited to run on GPUs, but newly developed algorithms usually require the heavily optimized DNN routines to work effi…