Compacting Deep Neural Networks for Internet of Things: Methods and Applications
arXiv:2103.11083 · doi:10.1109/JIOT.2021.3063497
Abstract
Deep Neural Networks (DNNs) have shown great success in completing complex tasks. However, DNNs inevitably bring high computational cost and storage consumption due to the complexity of hierarchical structures, thereby hindering their wide deployment in Internet-of-Things (IoT) devices, which have limited computational capability and storage capacity. Therefore, it is a necessity to investigate the technologies to compact DNNs. Despite tremendous advances in compacting DNNs, few surveys summarize compacting-DNNs technologies, especially for IoT applications. Hence, this paper presents a comprehensive study on compacting-DNNs technologies. We categorize compacting-DNNs technologies into three major types: 1) network model compression, 2) Knowledge Distillation (KD), 3) modification of network structures. We also elaborate on the diversity of these approaches and make side-by-side comparisons. Moreover, we discuss the applications of compacted DNNs in various IoT applications and outline future directions.
25 pages, 11 figures
References in corpus (28)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Learning in Neural Networks: An Overview
- Distilling the Knowledge in a Neural Network
- FitNets: Hints for Thin Deep Nets
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Deep Learning with Limited Numerical Precision
- Compressing Deep Convolutional Networks using Vector Quantization
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Pruning Filters for Efficient ConvNets
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- Dynamic Network Surgery for Efficient DNNs
- Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons
- ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks
- DeepIoT: Compressing Deep Neural Network Structures for Sensing Systems with a Compressor-Critic Framework
- Graph-based Knowledge Distillation by Multi-head Attention Network
- Weightless: Lossy Weight Encoding For Deep Neural Network Compression
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- Residual Knowledge Distillation
- Knowledge Flow: Improve Upon Your Teachers
- Knowledge Projection for Deep Neural Networks
- SCAN: A Scalable Neural Networks Framework Towards Compact and Efficient Models
- SEP-Nets: Small and Effective Pattern Networks
- LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuning
- Filter Grafting for Deep Neural Networks
- Graph Representation Learning via Multi-task Knowledge Distillation
- RTN: Reparameterized Ternary Network