Bringing AI To Edge: From Deep Learning's Perspective
arXiv:2011.14808 · doi:10.1016/j.neucom.2021.04.141
Abstract
Edge computing and artificial intelligence (AI), especially deep learning for nowadays, are gradually intersecting to build a novel system, called edge intelligence. However, the development of edge intelligence systems encounters some challenges, and one of these challenges is the \textit{computational gap} between computation-intensive deep learning algorithms and less-capable edge systems. Due to the computational gap, many edge intelligence systems cannot meet the expected performance requirements. To bridge the gap, a plethora of deep learning techniques and optimization methods are proposed in the past years: light-weight deep learning models, network compression, and efficient neural architecture search. Although some reviews or surveys have partially covered this large body of literature, we lack a systematic and comprehensive review to discuss all aspects of these deep learning techniques which are critical for edge intelligence implementation. As various and diverse methods which are applicable to edge systems are proposed intensively, a holistic review would enable edge computing engineers and community to know the state-of-the-art deep learning techniques which are instrumental for edge intelligence and to facilitate the development of edge intelligence systems. This paper surveys the representative and latest deep learning techniques that are useful for edge intelligence systems, including hand-crafted models, model compression, hardware-aware neural architecture search and adaptive deep learning models. Finally, based on observations and simple experiments we conducted, we discuss some future directions.
References in corpus (23)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Neural Architecture Search with Reinforcement Learning
- Mobile Edge Computing: A Survey on Architecture and Computation Offloading
- FitNets: Hints for Thin Deep Nets
- Deep Learning with Limited Numerical Precision
- Trained Ternary Quantization
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Towards Accurate Binary Convolutional Neural Network
- Channel Pruning for Accelerating Very Deep Neural Networks
- Light-Head R-CNN: In Defense of Two-Stage Object Detector
- Slimmable Neural Networks
- Adaptive Multi-Teacher Multi-level Knowledge Distillation
- Lite Transformer with Long-Short Range Attention
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- SpArSe: Sparse Architecture Search for CNNs on Resource-Constrained Microcontrollers
- Proving the Lottery Ticket Hypothesis: Pruning is All You Need
- ZynqNet: An FPGA-Accelerated Embedded Convolutional Neural Network
- DeepIoT: Compressing Deep Neural Network Structures for Sensing Systems with a Compressor-Critic Framework
- Defensive Quantization: When Efficiency Meets Robustness
- APQ: Joint Search for Network Architecture, Pruning and Quantization Policy
- Partial Order Pruning: for Best Speed/Accuracy Trade-off in Neural Architecture Search
- Enabling On-Device CNN Training by Self-Supervised Instance Filtering and Error Map Pruning
Cited by in corpus (4)
- Synthetic data generation method for data-free knowledge distillation in regression neural networks
- No Free Lunch: Balancing Learning and Exploitation at the Network Edge
- HSCoNAS: Hardware-Software Co-Design of Efficient DNNs via Neural Architecture Search
- Toward Compact Parameter Representations for Architecture-Agnostic Neural Network Compression