230 citations · 502 across the 5 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2021★ 13 cited
VS-Quant: Per-vector Scaled Quantization for Accurate Low-Precision Neural Network Inference
Steve Dai, Rangharajan Venkatesan, Haoxing Ren +3
Quantization enables efficient acceleration of deep neural networks by reducing model memory footprint and exploiting low-cost integer math hardware units. Quantization maps floati…
cs.LG2017★ 230 cited
Exploring the Regularity of Sparse Structure in Convolutional Neural Networks
Huizi Mao, Song Han, Jeff Pool +4
Sparsity helps reduce the computational complexity of deep neural networks by skipping zeros. Taking advantage of sparsity is listed as a high priority in next generation DNN accel…