HPTQ: Hardware-Friendly Post Training Quantization
arXiv:2109.09113
Abstract
Neural network quantization enables the deployment of models on edge devices. An essential requirement for their hardware efficiency is that the quantizers are hardware-friendly: uniform, symmetric, and with power-of-two thresholds. To the best of our knowledge, current post-training quantization methods do not support all of these constraints simultaneously. In this work, we introduce a hardware-friendly post training quantization (HPTQ) framework, which addresses this problem by synergistically combining several known quantization methods. We perform a large-scale study on four tasks: classification, object detection, semantic segmentation and pose estimation over a wide variety of network architectures. Our extensive experiments show that competitive results can be obtained under hardware-friendly constraints.
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Searching for Activation Functions
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Improving Neural Network Quantization without Retraining using Outlier Channel Splitting
- Simple and Lightweight Human Pose Estimation
- Fighting Quantization Bias With Bias
- HAWQV3: Dyadic Neural Network Quantization
- A White Paper on Neural Network Quantization