FATNN: Fast and Accurate Ternary Neural Networks
arXiv:2008.05101
Abstract
Ternary Neural Networks (TNNs) have received much attention due to being potentially orders of magnitude faster in inference, as well as more power efficient, than full-precision counterparts. However, 2 bits are required to encode the ternary representation with only 3 quantization levels leveraged. As a result, conventional TNNs have similar memory consumption and speed compared with the standard 2-bit models, but have worse representational capability. Moreover, there is still a significant gap in accuracy between TNNs and full-precision networks, hampering their deployment to real applications. To tackle these two challenges, in this work, we first show that, under some mild constraints, computational complexity of the ternary inner product can be reduced by a factor of 2. Second, to mitigate the performance gap, we elaborately design an implementation-dependent ternary quantization algorithm. The proposed framework is termed Fast and Accurate Ternary Neural Networks (FATNN). Experiments on image classification demonstrate that our FATNN surpasses the state-of-the-arts by a significant margin in accuracy. More importantly, speedup evaluation compared with various precisions is analyzed on several platforms, which serves as a strong benchmark for further research.
Accepted to Proc. Int. Conf. Computer Vision, ICCV 2021
References in corpus (13)
- Neural Architecture Search with Reinforcement Learning
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- DARTS: Differentiable Architecture Search
- Binarized Neural Networks
- Ternary Weight Networks
- Trained Ternary Quantization
- PACT: Parameterized Clipping Activation for Quantized Neural Networks
- Pruning Filters for Efficient ConvNets
- Loss-aware Weight Quantization of Deep Networks
- ProxQuant: Quantized Neural Networks via Proximal Operators
- Training Competitive Binary Neural Networks from Scratch
- Bi-Real Net: Enhancing the Performance of 1-bit CNNs With Improved Representational Capability and Advanced Training Algorithm
- Learning to Train a Binary Neural Network