Streamlined Deployment for Quantized Neural Networks
arXiv:1709.04060
Abstract
Running Deep Neural Network (DNN) models on devices with limited computational capability is a challenge due to large compute and memory requirements. Quantized Neural Networks (QNNs) have emerged as a potential solution to this problem, promising to offer most of the DNN accuracy benefits with much lower computational cost. However, harvesting these benefits on existing mobile CPUs is a challenge since operations on highly quantized datatypes are not natively supported in most instruction set architectures (ISAs). In this work, we first describe a streamlining flow to convert all QNN inference operations to integer ones. Afterwards, we provide techniques based on processing one bit position at a time (bit-serial) to show how QNNs can be efficiently deployed using common bitwise operations. We demonstrate the potential of QNNs on mobile CPUs with microbenchmarks and on a quantized AlexNet, which is 3.5x faster than an optimized 8-bit baseline. Our bit-serial matrix multiplication library is available on GitHub at https://git.io/vhshn
Presented at the International Workshop on Highly Efficient Neural Networks Design (HENND) co-located with CASES'17
References in corpus (3)
Cited by in corpus (8)
- Applications and Techniques for Fast Machine Learning in Science
- FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural Networks
- Memory-Driven Mixed Low Precision Quantization For Enabling Deep Network Inference On Microcontrollers
- Optimizing Bit-Serial Matrix Multiplication for Reconfigurable Computing
- Benchmarking Quantized Neural Networks on FPGAs with FINN
- BISMO: A Scalable Bit-Serial Matrix Multiplication Overlay for Reconfigurable Computing
- Quantized Neural Network Inference with Precision Batching
- Resource-Efficient Speech Mask Estimation for Multi-Channel Speech Enhancement