Faster Convolution Inference Through Using Pre-Calculated Lookup Tables
arXiv:2104.01681
Abstract
Low-cardinality activations permit an algorithm based on fetching the inference values from pre-calculated lookup tables instead of calculating them every time. This algorithm can have extensions, some of which offer abilities beyond those of the currently used algorithms. It also allows for a simpler and more effective CNN-specialized hardware.
11 pages, 7 figures
References in corpus (7)
- Trained Ternary Quantization
- Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition
- Analysis and Optimization of Convolutional Neural Network Architectures
- EasyQuant: Post-training Quantization via Scale Optimization
- Acceleration of Convolutional Neural Network Using FFT-Based Split Convolutions
- Towards Unified INT8 Training for Convolutional Neural Network
- Precision Scaling of Neural Networks for Efficient Audio Processing