Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead
arXiv:2012.11233 · doi:10.1109/ACCESS.2020.3039858
Abstract
Currently, Machine Learning (ML) is becoming ubiquitous in everyday life. Deep Learning (DL) is already present in many applications ranging from computer vision for medicine to autonomous driving of modern cars as well as other sectors in security, healthcare, and finance. However, to achieve impressive performance, these algorithms employ very deep networks, requiring a significant computational power, both during the training and inference time. A single inference of a DL model may require billions of multiply-and-accumulated operations, making the DL extremely compute- and energy-hungry. In a scenario where several sophisticated algorithms need to be executed with limited energy and low latency, the need for cost-effective hardware platforms capable of implementing energy-efficient DL execution arises. This paper first introduces the key properties of two brain-inspired models like Deep Neural Network (DNN), and Spiking Neural Network (SNN), and then analyzes techniques to produce efficient and high-performance designs. This work summarizes and compares the works for four leading platforms for the execution of algorithms such as CPU, GPU, FPGA and ASIC describing the main solutions of the state-of-the-art, giving much prominence to the last two solutions since they offer greater design flexibility and bear the potential of high energy-efficiency, especially for the inference process. In addition to hardware solutions, this paper discusses some of the important security issues that these DNN and SNN models may have during their execution, and offers a comprehensive section on benchmarking, explaining how to assess the quality of different networks and hardware systems designed for them.
Accepted for publication in IEEE Access
References in corpus (23)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Explaining and Harnessing Adversarial Examples
- ADADELTA: An Adaptive Learning Rate Method
- FitNets: Hints for Thin Deep Nets
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- cuDNN: Efficient Primitives for Deep Learning
- Compressing Deep Convolutional Networks using Vector Quantization
- Stealing Machine Learning Models via Prediction APIs
- Certified Adversarial Robustness via Randomized Smoothing
- Fast is better than free: Revisiting adversarial training
- A Survey of Neuromorphic Computing and Neural Networks in Hardware
- Robust Machine Learning Systems: Challenges, Current Trends, Perspectives, and the Road Ahead
- On Neural Architecture Search for Resource-Constrained Hardware Platforms
- Augment your batch: better training with larger batches
- Enabling Deep Spiking Neural Networks with Hybrid Conversion and Spike Timing Dependent Backpropagation
- CapsAttacks: Robust and Imperceptible Adversarial Attacks on Capsule Networks
- RED-Attack: Resource Efficient Decision based Attack for Machine Learning
- Closing the Accuracy Gap in an Event-Based Visual Recognition Task
- CapStore: Energy-Efficient Design and Management of the On-Chip Memory for CapsuleNet Inference Accelerators
- A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity