1 paper
Weihan Chen, Peisong Wang, Jian Cheng
Quantization is a widely used technique to compress and accelerate deep neural networks. However, conventional quantization methods use the same bit-width for all (or most of) the…