HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks
arXiv:1911.03852
Abstract
Quantization is an effective method for reducing memory footprint and inference time of Neural Networks, e.g., for efficient inference in the cloud, especially at the edge. However, ultra low precision quantization could lead to significant degradation in model generalization. A promising method to address this is to perform mixed-precision quantization, where more sensitive layers are kept at higher precision. However, the search space for a mixed-precision quantization is exponential in the number of layers. Recent work has proposed HAWQ, a novel Hessian based framework, with the aim of reducing this exponential search space by using second-order information. While promising, this prior work has three major limitations: (i) HAWQV1 only uses the top Hessian eigenvalue as a measure of sensitivity and do not consider the rest of the Hessian spectrum; (ii) HAWQV1 approach only provides relative sensitivity of different layers and therefore requires a manual selection of the mixed-precision setting; and (iii) HAWQV1 does not consider mixed-precision activation quantization. Here, we present HAWQV2 which addresses these shortcomings. For (i), we perform a theoretical analysis showing that a better sensitivity metric is to compute the average of all of the Hessian eigenvalues. For (ii), we develop a Pareto frontier based method for selecting the exact bit precision of different layers without any manual selection. For (iii), we extend the Hessian analysis to mixed-precision activation quantization. We have found this to be very beneficial for object detection. We show that HAWQV2 achieves new state-of-the-art results for a wide range of tasks.
References in corpus (1)
Cited by in corpus (21)
- Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors
- Bayesian Bits: Unifying Quantization and Pruning
- PyHessian: Neural Networks Through the Lens of the Hessian
- APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
- ZeroQ: A Novel Zero Shot Quantization Framework
- I-BERT: Integer-only BERT Quantization
- A Statistical Framework for Low-bitwidth Training of Deep Neural Networks
- ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training
- Channel-wise Hessian Aware trace-Weighted Quantization of Neural Networks
- Sharpness-aware Quantization for Deep Neural Networks
- Effective and Fast: A Novel Sequential Single Path Search for Mixed-Precision Quantization
- Differentiable Joint Pruning and Quantization for Hardware Efficiency
- MetaMix: Meta-state Precision Searcher for Mixed-precision Activation Quantization
- Post-Training Quantization for Vision Transformer
- Generative Design of Hardware-aware DNNs
- Exploring Neural Networks Quantization via Layer-Wise Quantization Analysis
- GradFreeBits: Gradient Free Bit Allocation for Dynamic Low Precision Neural Networks
- Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization
- PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
- Mixed Precision DNN Qunatization for Overlapped Speech Separation and Recognition
- RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple Precisions