A Survey of FPGA-Based Neural Network Accelerator
arXiv:1712.08934
Abstract
Recent researches on neural network have shown significant advantage in machine learning over traditional algorithms based on handcrafted features and models. Neural network is now widely adopted in regions like image, speech and video recognition. But the high computation and storage complexity of neural network inference poses great difficulty on its application. CPU platforms are hard to offer enough computation capacity. GPU platforms are the first choice for neural network process because of its high computation capacity and easy to use development frameworks. On the other hand, FPGA-based neural network inference accelerator is becoming a research topic. With specifically designed hardware, FPGA is the next possible solution to surpass GPU in speed and energy efficiency. Various FPGA-based accelerator designs have been proposed with software and hardware optimization techniques to achieve high speed and energy efficiency. In this paper, we give an overview of previous work on neural network inference accelerators based on FPGA and summarize the main techniques used. An investigation from software to hardware, from circuit level to system level is carried out to complete analysis of FPGA-based neural network inference accelerator design and serves as a guide to future work.
References in corpus (15)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
- SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Empirical Evaluation of Rectified Activations in Convolutional Network
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Ternary Weight Networks
- Trained Ternary Quantization
- Compressing Neural Networks with the Hashing Trick
- SkipNet: Learning Dynamic Routing in Convolutional Networks
- An OpenCL(TM) Deep Learning Accelerator on Arria 10
- Design Flow of Accelerating Hybrid Extremely Low Bit-width Neural Network in Embedded FPGA
- A GPU-Outperforming FPGA Accelerator Architecture for Binary Convolutional Neural Networks
Cited by in corpus (14)
- AddNet: Deep Neural Networks Using FPGA-Optimized Multipliers
- Pruning and Quantization for Deep Neural Network Acceleration: A Survey
- Mixed Precision DNNs: All you need is a good parametrization
- An Experimental Study of Reduced-Voltage Operation in Modern FPGAs for Neural Network Acceleration
- A Survey of FPGA Based Deep Learning Accelerators: Challenges and Opportunities
- PUMA: A Programmable Ultra-efficient Memristor-based Accelerator for Machine Learning Inference
- Evaluating Built-in ECC of FPGA on-chip Memories for the Mitigation of Undervolting Faults
- Intermediate Deep Feature Compression: the Next Battlefield of Intelligent Sensing
- A Generalized Zero-Shot Quantization of Deep Convolutional Neural Networks via Learned Weights Statistics
- An Efficient Hardware Accelerator for Structured Sparse Convolutional Neural Networks on FPGAs
- Dynamically Reconfigurable Variable-precision Sparse-Dense Matrix Acceleration in Tensorflow Lite
- A Low-Cost Neural ODE with Depthwise Separable Convolution for Edge Domain Adaptation on FPGAs
- On the Effects of Quantisation on Model Uncertainty in Bayesian Neural Networks
- Hardware Synthesis of State-Space Equations; Application to FPGA Implementation of Shallow and Deep Neural Networks