Quantized Convolutional Neural Networks Through the Lens of Partial Differential Equations
arXiv:2109.00095 · doi:10.1007/s40687-022-00354-y
Abstract
Quantization of Convolutional Neural Networks (CNNs) is a common approach to ease the computational burden involved in the deployment of CNNs, especially on low-resource edge devices. However, fixed-point arithmetic is not natural to the type of computations involved in neural networks. In this work, we explore ways to improve quantized CNNs using PDE-based perspective and analysis. First, we harness the total variation (TV) approach to apply edge-aware smoothing to the feature maps throughout the network. This aims to reduce outliers in the distribution of values and promote piece-wise constant maps, which are more suitable for quantization. Secondly, we consider symmetric and stable variants of common CNNs for image classification, and Graph Convolutional Networks (GCNs) for graph node-classification. We demonstrate through several experiments that the property of forward stability preserves the action of a network under different quantization rates. As a result, stable quantized networks behave similarly to their non-quantized counterparts even though they rely on fewer parameters. We also find that at times, stability even aids in improving accuracy. These properties are of particular interest for sensitive, resource-constrained, low-power or real-time applications like autonomous driving.
References in corpus (12)
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Knowledge Distillation: A Survey
- Language Models are Few-Shot Learners
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Simple and Deep Graph Convolutional Networks
- Improving Neural Network Quantization without Retraining using Outlier Channel Splitting
- ANODE: Unconditionally Accurate Memory-Efficient Gradients for Neural ODEs
- IMEXnet: A Forward Stable Deep Neural Network
- PDE-GCN: Novel Architectures for Graph Neural Networks Motivated by Partial Differential Equations
- LeanConvNets: Low-cost Yet Effective Convolutional Neural Networks
- Weight Normalization based Quantization for Deep Neural Network Compression
- GradFreeBits: Gradient Free Bit Allocation for Dynamic Low Precision Neural Networks