Toolflows for Mapping Convolutional Neural Networks on FPGAs: A Survey and Future Directions
arXiv:1803.05900
Abstract
In the past decade, Convolutional Neural Networks (CNNs) have demonstrated state-of-the-art performance in various Artificial Intelligence tasks. To accelerate the experimentation and development of CNNs, several software frameworks have been released, primarily targeting power-hungry CPUs and GPUs. In this context, reconfigurable hardware in the form of FPGAs constitutes a potential alternative platform that can be integrated in the existing deep learning ecosystem to provide a tunable balance between performance, power consumption and programmability. In this paper, a survey of the existing CNN-to-FPGA toolflows is presented, comprising a comparative study of their key characteristics which include the supported applications, architectural choices, design space exploration methods and achieved performance. Moreover, major challenges and objectives introduced by the latest trends in CNN algorithmic research are identified and presented. Finally, a uniform evaluation methodology is proposed, aiming at the comprehensive, complete and in-depth evaluation of CNN-to-FPGA toolflows.
Accepted for publication at the ACM Computing Surveys (CSUR) journal, 2018
References in corpus (12)
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Deep Learning with Limited Numerical Precision
- Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
- Binarized Neural Networks
- Trained Ternary Quantization
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights
- Hardware-oriented Approximation of Convolutional Neural Networks
- FPGA-Based CNN Inference Accelerator Synthesized from Multi-Threaded C Software
- Compiling Deep Learning Models for Custom Hardware Accelerators
- Scaling Binarized Neural Networks on Reconfigurable Logic