1 paper
Shoukai Xu, Haokun Li, Bohan Zhuang +4
Neural network quantization is an effective way to compress deep models and improve their execution latency and energy efficiency, so that they can be deployed on mobile or embedde…